An LLM that observes a real browser, plans one action, and loops until the task is done.
An agent that drives a real browser — clicking, typing, reading the page — so it can use software that has no API.
It unlocks the enormous amount of work that only exists behind a web UI, and it's the sharpest example of why containment matters.
A browser agent is an LLM given a real browser instead of a fixed script: it perceives each page as an accessibility tree plus a screenshot distilled into indexed interactive elements, plans a single next action, and executes it through a driver like Playwright or the Chrome DevTools Protocol. After every action the page changes, the agent re-observes and verifies, and it repeats this observe-plan-act loop, self-correcting on failure, until the goal is met. Frameworks such as browser-use, OpenAI's Operator, and Anthropic's computer use all run this same perception-action loop.
A browser agent turns a plain-language goal into real web actions. It observes the page as an accessibility tree plus a screenshot reduced to clickable elements, has the LLM choose one action, executes it with Playwright or CDP, then re-observes and verifies before looping again. The hard parts are dynamic DOMs, iframes, CAPTCHAs, and latency, so production systems add human approval for irreversible actions and defend against prompt injection from untrusted page text.
Browser Use Agent Tutorial: Automate Web Tasks with AI — Browser Use, 4:04