Single-pass CNN detector: grid cells predict boxes, NMS removes duplicates
Instead of scanning an image many times looking for objects, cut it into a grid and have every cell answer once: 'is something here, what, and where exactly?'
It's the reason object detection runs in real time on a phone or a camera feed, and the standard answer for any 'find the things' problem.
YOLO ('You Only Look Once') is a fast, single-stage object detector. Instead of the slow two-stage approach of proposing regions and then classifying them, a CNN backbone processes the whole image in one forward pass and outputs a grid whose cells each predict bounding boxes, an objectness score, and class probabilities. Because that yields many overlapping boxes, Non-Max Suppression keeps the best per object, giving real-time detection.
YOLO is a real-time object detector that looks at the whole image once. A single CNN forward pass produces an S-by-S grid; each cell predicts a few bounding boxes with an objectness score and class probabilities. That gives thousands of overlapping boxes, so Non-Max Suppression keeps the highest-confidence box per object and drops the rest. Being one-stage — no separate region-proposal step — is what makes it fast enough for video, unlike two-stage detectors like R-CNN.
How YOLO works for object detection | Computer Vision — AI Sciences, 5:05