All concepts

Support Vector Machine

Find the decision boundary with the largest margin from the nearest examples.

Classical ML · Intermediate · ~8 min

In plain English

Draw the boundary between two groups so that the empty corridor around it is as wide as possible. Only the points touching the corridor edges matter.

Why it's worth your time

It's the cleanest illustration of margin and the kernel trick — and still excellent on medium-sized, high-dimensional data like text.

If you remember three things

  • Only the support vectors define the boundary
  • C trades margin width against training errors
  • Kernels give non-linear boundaries without building the features

Overview

A classifier that finds the decision boundary with the maximum margin — the widest gap to the nearest points of each class. Only those nearest 'support vectors' define the boundary, and the kernel trick lets it draw nonlinear boundaries efficiently.

How it works

  1. Start: Labeled Points Training examples occupy a feature space with class labels.
  2. Labeled Points -> Margin The SVM searches for a boundary that leaves the widest gap to both classes.
  3. Margin -> Support Vectors Only the nearest boundary-defining points matter for the final classifier.
  4. Support Vectors -> Kernel Trick A kernel can create nonlinear boundaries without explicitly materializing high-dimensional features.
  5. Kernel Trick -> Classify A new point is assigned based on which side of the maximum-margin boundary it falls.

In an interview

An SVM finds the maximum-margin separating hyperplane, minimizing ½‖w‖² + C·Σ hinge_loss. Only the closest points, the support vectors, determine it. The C parameter trades margin width against violations, and kernels (RBF, polynomial) let it separate non-linearly-separable data without explicitly building high-dimensional features.

Production defaults

Scale
mandatory. An unscaled SVM is a broken SVM
Search
C and γ on a log grid (1e-3 … 1e3). They interact, so search jointly, not one at a time
Kernel
linear for text and wide sparse data; RBF for dense tabular

What breaks

  • Training never finishes — SVM training scales roughly quadratically. Above ~100k rows use LinearSVC/SGDClassifier instead.
  • Everything classified as one class — γ far too high — the RBF kernel has collapsed to memorizing points. Drop it by orders of magnitude.

Watch it explained

Support Vector Machine (SVM) in 2 minutes — Visually Explained, 2:19

Related