Concept library

208 concepts across AI/ML Engineering, Data Engineering and Data Analytics — each with an animation, an interview answer and a production recipe.

ML Foundations

  • Linear Regression — Fit the straight line that makes the total squared distance to your data as small as possible.
  • Gradient Descent — Walk downhill on the loss surface by repeatedly stepping opposite the gradient.
  • Logistic Regression — Turn a linear score into a probability with the sigmoid, then threshold it to classify.
  • Bias–Variance Tradeoff — Too simple underfits (high bias); too flexible overfits (high variance). Generalization lives in between.
  • Cross-Validation — Rotate which fold is held out for validation, then average — a stable estimate from limited data.
  • One-Hot Encoding — Represent a category as a sparse vector with exactly one active position.
  • Regularization — Add a penalty to discourage overly complex models that memorize noise.
  • Feature Scaling — Put features on comparable numeric ranges so distances and gradients behave.
  • Optimizers — Convert gradients into parameter updates using momentum, adaptive scaling, and schedules.
  • Vanishing Gradient — Gradients shrink as they flow backward, so early layers barely learn.

Classical ML

  • Decision Tree — Recursively split the feature space along the questions that most reduce impurity.
  • K-Means Clustering — Alternate between assigning points to the nearest centroid and moving each centroid to its cluster's mean.
  • Principal Component Analysis — Rotate the axes to line up with the directions of greatest variance, then keep the top few.
  • Random Forest — Average many decorrelated decision trees to cut variance and boost accuracy.
  • Gradient Boosting / XGBoost — Add shallow trees one at a time, each correcting the residual errors of the last.
  • K-Nearest Neighbors — Predict by asking the closest labeled examples to vote.
  • Naive Bayes — Use Bayes rule with a simplifying independence assumption for fast classification.
  • Support Vector Machine — Find the decision boundary with the largest margin from the nearest examples.
  • Ensemble Methods — Combine many models so their errors cancel or specialize.
  • t-SNE — Visualize high-dimensional points by preserving local neighborhoods in 2D.

Deep Learning

  • Neural Network Forward Pass — Data flows layer by layer: linear combination, non-linear activation, repeat — producing a prediction.
  • Backpropagation — Apply the chain rule backward through the network to get every weight's gradient in one efficient pass.
  • Activation Functions — The non-linear bend that lets stacked layers model complex functions — sigmoid, tanh, ReLU, GELU.
  • CNN Convolution — A small kernel slides over the image, computing dot products to detect local patterns.
  • CLIP Alignment — Train image and text encoders so matching captions land near matching images.
  • Recurrent Neural Network — Process sequences one step at a time while carrying hidden state forward.
  • Vision Transformer (ViT) — Treat an image as a sequence of patches and run a transformer over them.
  • Diffusion Denoising — Learn to reverse a gradual noising process to create samples.
  • Distributed Training — Split training work across many GPUs or machines while keeping gradients synchronized.
  • YOLO Object Detection — Single-pass CNN detector: grid cells predict boxes, NMS removes duplicates

Transformers & LLMs

  • BPE Tokenization — Start from characters and greedily merge the most frequent adjacent pair, building a subword vocabulary.
  • Self-Attention — Each token builds a query, matches it against every key, and pulls a weighted blend of values.
  • Transformer Block — Attention + feed-forward, each wrapped in residual connections and normalization.
  • KV Cache — Cache past keys and values so each new token costs O(n) instead of O(n²).
  • LoRA / QLoRA — Freeze the base model and train tiny low-rank adapter matrices instead of all the weights.
  • Token Sampling (Temperature, Top-k, Top-p) — How an LLM turns logits into a chosen next token — softmax, temperature, then top-k / top-p.
  • RLHF & DPO (Alignment) — Turn human preferences into model behavior — via a reward model + RL (RLHF) or directly (DPO).
  • Chain-of-Thought & Self-Consistency — Make the model reason step by step, then vote across several chains for a reliable answer.
  • Mixture of Experts (MoE) — A router sends each token to a few expert FFNs — huge capacity at a fraction of the compute.
  • Transformer Architecture — The full encoder–decoder: embed, attend, add & norm, cross-attend, then predict the next token.
  • Multi-Head Attention — Run several attention patterns in parallel so tokens can attend for different reasons.
  • Attention Masking — Block illegal attention links so tokens only see what they are allowed to see.
  • Cross Attention — Let one sequence query information from another sequence.
  • Flash Attention — Compute exact attention while reducing expensive GPU memory traffic.
  • Grouped Query Attention — Share key/value heads across groups of query heads to shrink the KV cache.
  • Layer Norm — Normalize each token representation to stabilize deep network training.
  • Positional Encoding — Inject token order because attention alone is permutation-invariant.
  • RoPE Positional Encoding — Encode relative positions by rotating query and key vectors.
  • Context Window — The maximum tokens the model can read and generate in one request.
  • GPT Generation Loop — Generate one token at a time, feeding each chosen token back into the model.
  • Speculative Decoding — Use a small draft model to propose tokens that a large model verifies in batches.
  • Structured Output — Constrain model responses to JSON, schemas, or tool-call shapes.
  • Zero-shot vs One-shot vs Few-shot — Control task behavior by giving no examples, one example, or several examples in the prompt.
  • Instruction Tuning — Fine-tune a pretrained model on instruction-response examples so it follows tasks.
  • Quantization — Store weights or activations in fewer bits to reduce memory and speed serving.

RAG & Retrieval

  • RAG Pipeline — Retrieve relevant chunks for a query, stuff them into the prompt, and let the LLM answer with citations.
  • Hybrid Retrieval + Reranking — Combine keyword (BM25) and vector search, then rerank with a cross-encoder for precision.
  • Corrective RAG — Grade retrieved chunks; if they're weak, rewrite the query or fall back to web search before answering.
  • HNSW Vector Index — How graph-based approximate nearest-neighbour search finds close vectors in ~log(N) time.
  • How a Vector Database Works — Embed, index, and search: the anatomy of a vector database for RAG.
  • Cross-Encoder Reranking — Retrieve wide and cheap with a bi-encoder, then rerank narrow and precise with a cross-encoder.
  • Citations & Grounding — Tie generated claims back to retrieved source chunks so answers are auditable.
  • GraphRAG — Use entities and relationships to retrieve connected evidence, not just nearest chunks.
  • BM25 / TF-IDF — Score documents by lexical term overlap with saturation and rarity weighting.
  • Chunking Strategies — Split documents into retrieval units that are small enough to focus but large enough to make sense.
  • Semantic Cache — Reuse prior answers or retrieved contexts when a new query is semantically similar.
  • Knowledge Graphs (Neo4j) — Entities as nodes, relations as typed edges; query with Cypher, power GraphRAG

NLP & Embeddings

  • Embedding Vector Space — Map text to vectors where semantic similarity becomes geometric closeness.
  • Tokenization — Split text into model-readable tokens before embedding or generation.

Agentic AI

  • Agent Tool Calling — The model reasons, calls a tool, observes the result, and repeats until it can answer.
  • LangGraph Workflow — Model an agent as a stateful graph of nodes and edges with memory and checkpoints.
  • ReAct Agent — Alternate reasoning, tool actions, and observations until the task is solved.
  • Agent Memory — Store useful past information so an agent can personalize and continue tasks.
  • Agent Planning — Break a goal into ordered or conditional sub-tasks before acting.
  • Agent Reflection — Have the agent critique its own output or trajectory before finalizing.
  • Multi-Agent Systems — Coordinate multiple specialized agents to solve a task collaboratively.
  • MCP Protocol — Expose tools and resources to AI clients through a standardized protocol.
  • OpenAI Agents SDK — OpenAI's lightweight SDK for agents: an LLM loop with tools, handoffs, and guardrails.
  • Claude Agent SDK — Anthropic's agent harness — the gather → act → verify loop behind Claude Code, plus computer use.
  • Browser Agents — An LLM that observes a real browser, plans one action, and loops until the task is done.
  • LangChain — Framework to compose LLM apps: prompt | model | parser as LCEL Runnables

MLOps & LLMOps

  • LLMOps Monitoring — Track quality, latency, token cost, and drift — with evals and tracing — to keep LLM apps healthy.
  • Model Routing & Token Cost — Send easy queries to cheap models and hard ones to strong models to cut cost without losing quality.
  • Prompt Injection & Guardrails — Untrusted content can hijack an LLM agent — treat tool/retrieved text as data, never as commands.
  • Data Drift — Detect when production input distributions move away from training data.
  • Feature Store — A shared system for defining, computing, serving, and monitoring ML features.
  • Model Monitoring — Track model quality, latency, cost, drift, and failures after deployment.
  • Shadow Deployment — Run a new model on production traffic without serving its answers yet.
  • Training-Serving Skew — A mismatch between training-time data processing and production-time serving.
  • LangSmith — Trace, evaluate, and monitor LLM apps: nested spans, datasets, prod feedback
  • Langfuse — Open-source LLM observability: traces, scores, prompt mgmt, self-hostable

Model Evaluation

  • Confusion Matrix & Metrics — Precision, recall, and F1 all fall out of the 2×2 table of TP / FP / FN / TN.
  • ROC Curve & AUC — Sweep every threshold, plot TPR vs FPR — the area underneath (AUC) is threshold-free quality.
  • LLM as a Judge — Use a stronger model with a rubric to grade answers at scale.

Production AI Systems

  • Agentic RAG on AWS — Build an agentic RAG system on AWS with Bedrock Agents, Knowledge Bases, OpenSearch, and Guardrails.
  • Agentic RAG on GCP — Build an agentic RAG system on GCP with Cloud Run, Vertex AI Agent Builder, RAG Engine, Vector Search, and Gemini.

Applied ML

  • Multistage Ranking — Rank items in stages: broad retrieval, lightweight scoring, heavy reranking, and final business rules.
  • Two Tower — Embed users and items separately so nearest-neighbor search can retrieve recommendations quickly.

Python

  • Decorators — Wrap a function to add behavior—logging, timing, auth—without touching its code.
  • Generators — yield values lazily, one at a time—stream huge data in constant memory.
  • Async / Await — async/await overlaps I/O waits on one thread—concurrency without threads.
  • Context Managers — with-blocks acquire and always release a resource—cleanup guaranteed, even on error.
  • Type Hints — Annotate types so checkers catch bugs before runtime; frameworks read them too.
  • Comprehensions — Build lists, dicts, and sets in one readable expression from an iterable + filter.
  • Pydantic — Typed models validate and coerce untrusted input at the boundary, with clear errors.
  • FastAPI — Type-hinted async endpoints with auto validation and docs—the default for ML APIs.

SQL

  • SELECT & Filtering — Describe which rows and columns you want; SQL filters, sorts, then limits.
  • SQL Joins — Match rows across tables on a key; the join type decides which non-matches survive.
  • GROUP BY & Aggregates — GROUP BY buckets rows; aggregates summarize each bucket; HAVING filters the groups.
  • SQL Indexes — A B-tree index turns a slow table scan into a fast lookup, at the cost of writes.
  • Subqueries & CTEs — A query inside a query; a CTE (WITH …) names it for readable, reusable, recursive steps.
  • Window Functions — OVER() computes across related rows without collapsing them — ranks, running totals, lag.
  • Transactions & ACID — BEGIN…COMMIT groups writes into one atomic, ACID unit; ROLLBACK undoes it on failure.

Cloud Services

  • AWS S3 — Durable object storage in buckets — the cheap, near-infinite data lake behind ML on AWS
  • AWS SageMaker — AWS's managed ML platform: train on GPUs, version models, and serve them as endpoints
  • AWS Bedrock — Managed foundation models behind one API, with RAG, agents, and guardrails built in
  • GCP Cloud Storage — Object storage buckets on GCP — the tiered, durable data lake for Vertex and BigQuery
  • Vertex AI — Google Cloud's unified ML platform: training, serving, Vector Search, Gemini, and agents
  • BigQuery — Serverless columnar data warehouse — SQL over petabytes, storage split from compute

Maths

  • Mean, Median & Mode — Summarize a dataset's center with one number: mean, median, or mode.
  • Variance & Std Dev — Measure how far data spreads from its mean, in the data's own units.
  • Probability Basics — Outcomes, events, and the rules for combining chances with AND and OR.
  • Bayes' Theorem — Update a prior belief with evidence to get a posterior probability.
  • Distributions — How probability spreads across a variable's values: normal, uniform, binomial.
  • Expectation & Variance — E[X] is a random variable's long-run average; Var[X] is its spread.
  • Correlation & Covariance — Do two variables move together? Covariance gives direction, correlation strength.
  • Hypothesis Testing — Test a claim by asking how surprising your data is if nothing were going on.

Advanced Embeddings

  • Matryoshka Embeddings — Train one embedding whose first 64 dimensions are already a usable embedding — so you can truncate for speed and only pay full width when it matters.
  • Late Interaction (ColBERT) — Keep one vector per token instead of one per document, and score with MaxSim — cross-encoder quality at nearly bi-encoder speed.
  • Learned Sparse Retrieval (SPLADE) — Let the language model decide the term weights — a sparse vector over the vocabulary that expands documents with terms they never contained.
  • Fine-tuning Embeddings — Contrastive training on your own query–document pairs, with mined hard negatives — usually the single biggest retrieval win available.

Advanced Vector Search

  • Product Quantization (IVF-PQ) — Split a vector into subspaces, replace each chunk with a codebook id, and search a billion vectors in RAM you can actually afford.
  • Binary Quantization + Rescoring — Keep one bit per dimension, search with XOR and popcount, then rescore the survivors — 32× less memory for a couple of points of recall.
  • Filtered Vector Search — Combining a metadata filter with an ANN graph is where recall quietly collapses — and why pre-, post-, and in-filter search are different products.
  • DiskANN & Billion-Scale Search — Put the graph on SSD, keep a compressed copy in RAM to steer, and serve a billion vectors from one machine.

Advanced RAG

  • Contextual Retrieval — Prepend an LLM-written sentence of surrounding context to every chunk before you embed it — the cheapest large retrieval win there is.
  • RAPTOR Hierarchical Index — Cluster chunks, summarise each cluster, cluster the summaries, and index every level — so one retriever answers both detail and 'what is this about' questions.
  • HyDE & Query Expansion — Search with a hallucinated answer instead of the question — a fake document is closer to the real one than a query ever is.
  • Self-RAG & Reflective Retrieval — Let the model decide whether to retrieve, then grade its own evidence and its own answer — retrieval on demand instead of retrieval on reflex.
  • Agentic RAG Controller — Retrieval as a controlled loop, not a function call: route, decompose, retrieve per sub-question, grade, escalate, and stop on a budget.

Advanced LLM Systems

  • PagedAttention & Continuous Batching — Treat the KV cache like virtual memory — pages, not contiguous blocks — and the GPU serves several times more concurrent requests.
  • Prompt Caching — The model already computed the KV for your system prompt — put the stable part first, mark it, and stop paying for it every call.
  • Constrained Decoding — Mask the logits so only tokens that keep the output valid can be sampled — schema conformance by construction, not by retrying.
  • Test-Time Compute & Reasoning Models — Buy accuracy with inference tokens instead of training runs — think longer, sample more, verify, and pick.
  • Long-Context Scaling — How 4k-trained models came to read 200k tokens — interpolate the positions, and know that 'fits' is not 'attends'.

Advanced Agentic Systems

  • Mixture of Agents — Several models answer, a later layer reads all their answers and writes a better one — ensembling for language, in layers.
  • Agent Sandboxing — An agent that writes and runs code is remote code execution with a friendly interface — isolate it like you'd isolate an untrusted binary.
  • Deep Research Agents — Plan, fan out subagents with separate context windows, verify every claim against a source, then synthesise — the shape behind every 'deep research' product.

DE Foundations

  • OLTP vs OLAP — One database is built to change a single row fast; the other is built to read a billion rows fast. They are not the same machine.
  • Star Schema — One long, skinny table of things that happened, surrounded by short, wide tables describing the nouns involved.
  • Normalization vs Denormalization — Store each fact once so it can never disagree with itself — or copy it everywhere so nobody has to join.
  • Columnar Storage & Parquet — Store a table column by column and a query that wants two of forty columns reads two of forty columns.
  • Partitioning & Clustering — Put the data in folders named after the thing you filter on, so most queries never open most folders.
  • Lake, Warehouse & Lakehouse — Cheap files with no rules, a strict database with every rule, or files plus a transaction log that gives you both.
  • Batch vs Streaming — Wait and process a pile every hour, or process each record as it lands — the difference is what you're willing to pay for freshness.

Pipelines & Orchestration

  • ETL vs ELT — Transform before you load and you throw away what you didn't anticipate; load first and the warehouse does the work, but it stores your mess.
  • Orchestration & DAGs — Declare which task depends on which and let the scheduler decide what can run now, what must wait, and what to retry.
  • Idempotency & Backfills — Running the same job twice must leave the world exactly as it was after running it once — otherwise every retry corrupts your data.
  • dbt & the Transformation Layer — Every table is a SELECT statement in version control, and the tool works out the order to run them in.
  • Data Quality Tests — A pipeline that succeeds is not the same as a pipeline that is correct — so assert the things that must be true, and stop the run when they aren't.
  • Schema Evolution & Data Contracts — Upstream renamed a column on a Tuesday and nobody told you — a contract is how that becomes their build failure instead of your incident.
  • Slowly Changing Dimensions — A customer moves from London to Berlin. Should last year's orders now count as German? Your answer is the SCD type.

Streaming & CDC

  • Kafka & the Commit Log — Not a queue that hands out messages and forgets them — an append-only log that keeps them, and readers who remember their own place in it.
  • Change Data Capture — Read the database's own write-ahead log instead of asking it what changed — you get every change, including the deletes a query can't see.
  • Event Time & Watermarks — A phone goes into a tunnel and its 9:00 event arrives at 9:07 — a watermark is the system's declaration of how long it will wait before closing the 9:00 window.
  • Windowing & Streaming Aggregation — An unbounded stream has no end, so you can't sum it — you cut it into windows, and how you cut it decides what the number means.
  • Delivery Semantics — At-most-once loses records, at-least-once duplicates them, and exactly-once is at-least-once plus somewhere to deduplicate.

Distributed Processing

  • Spark's Execution Model — Nothing runs until you ask for a result — then the plan is cut into stages at every point data has to cross the network.
  • The Shuffle — Every row with the same key has to end up on the same machine, and getting it there means writing the whole dataset to disk and reading it back across the network.
  • Joins at Scale — If one side fits in memory, ship it to every machine and the join is free — otherwise both sides get shuffled, sorted, and merged.
  • Data Skew — 199 tasks finish in a minute and one runs for two hours — the cluster isn't slow, one key owns half your data.
  • Predicate Pushdown & Query Planning — The optimiser's whole job is to move your WHERE clause as close to the disk as it can get it.

Data Platform in Production

  • Data Observability — The pipeline was green all week. The table stopped updating on Tuesday. Nobody noticed until Friday.
  • Lineage & Impact Analysis — Two questions, one graph: where did this number come from, and what breaks if I change this column?
  • Warehouse Cost & Performance — The bill is bytes scanned times how often you scan them — and almost every surprise on it is a dashboard refreshing every five minutes over five years of data.
  • Governance & PII — Know which columns are personal, restrict who can read them, and be able to delete one person from a lake that was designed to be immutable.
  • Pipeline Failure & On-Call — The pipeline will fail. What decides whether that's an inconvenience or an incident is whether the failure stops before it reaches a published table.

Analytics Foundations

  • From Question to Metric — "Is the new checkout working?" is not a question a query can answer — turning it into one is most of the job.
  • Sampling & Selection Bias — Your survey says 92% of users love the redesign. It was shown in-app, to people still using the app.
  • Correlation vs Causation — Users who enable notifications retain twice as well — so should you force notifications on, or are you just describing people who already liked the product?
  • Simpson's Paradox — Version B wins on mobile, wins on desktop, and loses overall. Both facts are true.

Analytics SQL

  • Cohort Analysis — Group users by when they joined, then follow each group forward — the only way to tell a product that is improving from one that is just growing.
  • Funnel Analysis — Count how many people reach each step in order, and the biggest drop tells you where to spend next quarter.
  • Retention Curves — Every retention curve falls. The only question that matters is whether it flattens, and how high.
  • Running Totals & Moving Averages — Window functions let a row see its neighbours — the running total, last week's value, the customer's previous order — without collapsing the rows away.

Metrics & KPIs

  • Defining a Metric — Three teams report "active users" and get three numbers. None of them is wrong; there was never one definition.
  • Metric Trees — Revenue fell 8%. A metric tree turns that sentence into four candidate causes in ninety seconds.
  • DAU, MAU & Stickiness — DAU/MAU says what fraction of your monthly users show up on an average day — a number that is 0.6 for a messaging app and 0.05 for a tax product, and both can be healthy.
  • Unit Economics: LTV & CAC — If a customer is worth £180 and costs £60 to acquire, growth is an investment. Reverse those and growth is how you go bust faster.
  • Guardrail Metrics — Every metric can be moved by making the product worse — a guardrail is the number that catches you doing it.

Experimentation

  • Anatomy of an A/B Test — Split users at random, change one thing, and the difference you measure is caused by the change — that last clause is the whole point, and randomisation is what buys it.
  • p-values & Significance — A p-value is the probability of seeing a difference this big if the change did nothing — not the probability that the change did nothing.
  • Confidence Intervals — "+2.1%" is a guess. "+2.1%, somewhere between +0.4% and +3.8%" is a result you can make a decision with.
  • Power & Sample Size — Work out how many users you need before you start, or you'll spend three weeks proving nothing and call it a null result.
  • Sample Ratio Mismatch — You asked for a 50/50 split and got 50.4/49.6. On two million users that is not rounding — it means something is filtering your users, and the result is void.

Dashboards & Storytelling

  • Choosing the Right Chart — The chart type isn't a style choice — it's a claim about what kind of comparison you want the reader to make.
  • Dashboard Design — A dashboard with forty charts is a place people go to feel informed. A dashboard with six is a place they go to decide something.
  • Telling the Story — Lead with the answer. The analysis is your evidence, not your narrative arc.
  • The Semantic Layer — Define 'revenue' once, in one place, and every dashboard, notebook and AI assistant that asks for revenue gets the same number.

Agentic Engineering

  • Loop Engineering — Bound the reason → act → observe cycle: name every exit, cap every budget, and detect a loop that has stopped learning.
  • Context Engineering — Treat the window as a budget: decide what stays resident, compress the middle, and place what matters where the model still reads it.
  • Tool Design — The tool surface IS the agent's API: precise descriptions, strict schemas, errors that teach, and fewer tools than you think.
  • Memory Architecture — Working memory is in-window and volatile; long-term memory is written out and retrieved — the policy between them is the design.
  • Orchestration Patterns — Single agent by default; add an orchestrator, a handoff, or parallelism only when a specific constraint forces it.
  • Guardrails & Permissions — Scope tools per task, separate read from write, filter both boundaries, and design so the worst single call is survivable.
  • Evals for Agents — Score the path as well as the answer, grow your golden set from real incidents, and gate every change on a regression run.
  • Human-in-the-Loop Design — Gate what can't be undone, escalate on a threshold you chose deliberately, review async, and absorb corrections without restarting.
  • Observability & Tracing — One structured span per step, reassembled into a trace — then cost, latency and failures become things you can see instead of guess.