All concepts

Browser Agents

An LLM that observes a real browser, plans one action, and loops until the task is done.

Agentic AI · Advanced · ~8 min

In plain English

An agent that drives a real browser — clicking, typing, reading the page — so it can use software that has no API.

Why it's worth your time

It unlocks the enormous amount of work that only exists behind a web UI, and it's the sharpest example of why containment matters.

If you remember three things

  • The page is untrusted input — it can contain instructions
  • Slow and brittle compared to an API; use one if it exists
  • Every irreversible click needs a human gate

Overview

A browser agent is an LLM given a real browser instead of a fixed script: it perceives each page as an accessibility tree plus a screenshot distilled into indexed interactive elements, plans a single next action, and executes it through a driver like Playwright or the Chrome DevTools Protocol. After every action the page changes, the agent re-observes and verifies, and it repeats this observe-plan-act loop, self-correcting on failure, until the goal is met. Frameworks such as browser-use, OpenAI's Operator, and Anthropic's computer use all run this same perception-action loop.

How it works

  1. Start: Goal A plain-language task: book a flight, extract these fields.
  2. Goal -> Observe Accessibility tree + screenshot, distilled to indexed elements.
  3. Observe -> Plan action The LLM chooses exactly one next action.
  4. Plan action -> Act Playwright / CDP executes click, type, scroll, navigate.
  5. Act -> Safety gate Sandbox, human approval for irreversible actions, injection defense.
  6. Safety gate -> Verify + loop Did the state change as expected? Re-observe and self-correct.
  7. Verify + loop -> Done Goal confirmed or data returned.

In an interview

A browser agent turns a plain-language goal into real web actions. It observes the page as an accessibility tree plus a screenshot reduced to clickable elements, has the LLM choose one action, executes it with Playwright or CDP, then re-observes and verifies before looping again. The hard parts are dynamic DOMs, iframes, CAPTCHAs, and latency, so production systems add human approval for irreversible actions and defend against prompt injection from untrusted page text.

Production defaults

Prefer the API
always, if one exists. Browser automation is the fallback, not the default
Trust
page content is data, never commands. This is not optional
Gates
purchases, sends, deletes and credential entry stop for a human, every time
Robustness
target by accessible role/label rather than pixel coordinates where you can

What breaks

  • The agent followed instructions embedded in a page — Indirect prompt injection — the defining risk of this pattern. Contain side effects outside the model.
  • Breaks whenever the site changes — Inherent brittleness. Selector-based targeting and a small retry budget help; an API helps more.

Watch it explained

Browser Use Agent Tutorial: Automate Web Tasks with AI — Browser Use, 4:04

Related