Repository navigation
when the page changes under the agent
Every browser agent runs the same loop: look, decide, act. The decide step is the slow one, because a model is thinking, and a web page does not wait. Everything awkward about driving a browser with a model lives in that gap.
It is worth being precise about how long the gap is. The looking and the acting are fast: a snapshot comes back in about a hundredth of a second and a click takes about a second, both measured. The thinking in between is seconds. So the page has, routinely, several seconds in which to become a different page while the agent is holding a description of the old one.
- A consent or cookie wall appears, and now everything the snapshot described is behind it.
- Images finish loading and push the layout down. Nothing changed logically; every coordinate moved.
- A lazy list renders and the element that was third is now ninth.
- A single-page app swaps a view, and the elements the snapshot listed no longer exist.
- A toast or a modal opens over the thing the agent was about to click.
- A session expires, and the page is now a login form that looks nothing like the snapshot.
None of these is a bug in the page, and none of them is the model being careless. They are the normal behaviour of the modern web meeting a client that thinks for three seconds.
This is the practical asymmetry, and it decides how you should drive.
A selector resolved after the page moved either finds the right element or finds nothing. Finding nothing is an error the agent can see and react to.
A coordinate resolved after the page moved always finds something: whatever is now at that x and y. The action succeeds, the wrong thing happens, and nothing in the run reports a problem.
That is why the sensible order is a named tool with a selector first, coordinates only when no selector describes the target, and a picture only when the snapshot does not list the thing at all. It is not about precision; it is about which kind of failure you get. Why did the AI agent click the wrong thing is the same asymmetry seen from the symptom, and clicking by selector or by coordinates has it with the timings attached, including the case where the selector cannot work at all and coordinates are the only route.
Re-read immediately before acting, and act on the fresh read.
It sounds obvious and it is routinely skipped, because a snapshot feels like it costs something. It costs a hundredth of a second. The right habit is:
snapshot -> decide -> snapshot again -> act on the second one
The second read closes the gap to roughly the duration of one tool call instead of the duration of one model turn, which is one to two orders of magnitude smaller. It does not close the gap to zero, and nothing can.
A common half-fix is to insert a wait: pause two seconds after navigating, then carry on. That helps with the slowest case and leaves the others alone, because the thing you are waiting for is not time, it is a state.
The better instruction names the state: wait until the result list has rows, wait until the spinner is gone, wait until the button is enabled. The agent can check those by reading, which is free. A fixed wait is a guess that is too long when the page is fast and too short when it is slow.
Three situations, three different responses.
The action failed cleanly. The best case. Re-read and retry once, on the fresh description. If the second attempt also fails, stop: something structural changed and retrying is now the loop described in giving an AI browser agent a stopping condition.
The action succeeded but the result is wrong. Re-read and compare against what you expected the page to say. This is the case that needs a stated expectation in the task, because without one there is nothing to compare against.
The action was irreversible. There is no recovery here, which is why this class of action needs a confirmation step in front of it rather than a retry behind it. That argument is in keeping an AI browser agent out of destructive actions.
A session that expires mid-run turns the page into a login form, and every subsequent step operates on the wrong page while looking superficially normal. The agent fills fields, the fields exist, nothing errors.
The defence is a cheap assertion rather than cleverness: have the task state one thing that must be true of the page it is on, and check it after each navigation. "The header shows my account name." It costs one free read and it turns a silent wrong run into an early stop.
Why does my agent act on stale information? Because it looked, then thought for a few seconds, then acted. The page had those seconds. Re-read immediately before acting and the window shrinks to a single tool call.
Should I add a wait after navigating? Prefer waiting for a state over waiting for a duration. A fixed wait is too long on a fast page and too short on a slow one.
Why did the click succeed and do the wrong thing? You clicked a coordinate, and the layout shifted. Coordinates always hit something; selectors fail cleanly.
How do I detect that the session expired mid-run? Give the task one sentence that must be true of a logged-in page and check it after every navigation.
Does a faster model fix this? It narrows the gap, which helps, and it does not close it. The structural fix is the second read, not the faster turn.
See also: how to write a task an AI browser agent can follow, where the "what must be true afterwards" sentence comes from, and running one task across a list of sites, where one flaky page in a list of fifty is the thing that decides whether the run is usable.
- The two timings this page leans on, a read at about 0.01 s and a click at about 1 s, are measured and sourced on how long an AI browser agent takes per step. The list of things that move during the gap is from running this browser rather than from a study, and is offered as such.
The whole page reduces to one line: the gap between looking and acting is a model turn long, and you can make it a tool call long instead. Everything else here is a consequence.
- OpenAI Operator alternatives
- Open-source Operator-style agents
- Is OpenAI Operator still available?
- OpenAI Operator vs Claude computer use
- browser-use alternatives
- Choosing an AI browser agent
- Open-source AI browser agents
- Open-source computer-use agents
- What is an AI web agent?
- AI browser agents vs traditional scraping
- Cloud browser infrastructure for AI agents, explained
- Browserbase alternatives
- Firecrawl vs an AI browser agent
- Skyvern alternatives
- Stagehand vs browser-use
- Project Mariner is gone: what replaced it
- Manus alternatives
- Gemini computer use vs Claude computer use
- A stealth browser MCP, reviewed honestly by its own wiki
- AI browser vs AI browser agent: which one do you want?
- AI browser agent vs RPA: which one fits the job
- AI browser agent vs n8n, Zapier and Make
- Vercel agent-browser alternatives, compared honestly
- What is an agentic browser? Definition and the two kinds
- Open-source agentic browsers: the three layers, compared
- Choosing an MCP server for browser automation: four axes
- Stealth MCP servers compared: Camoufox, nodriver, Patchright
- Playwright MCP alternatives, and the three you don't need
- Autonomous browser agents: the four rungs of autonomy
- What is actually free in the AI browser agent stack
- browser-use on GitHub: what the repo actually gives you
- Playwright MCP vs Chrome DevTools MCP: different jobs
- How to choose among MCP servers: a map by category
- Which MCP servers are worth adding to Claude Code
- MCP on GitHub: finding servers and judging them fast
- MCP vs an API: the decision, and what the wrapper costs
- MCP alternatives: when the protocol is the wrong shape
- Why does my AI agent get blocked?
- The timing signal AI agents give off
- Agent retry loops trip rate limits, not fingerprints
- Claude computer use detected as a bot
- browser-use getting blocked: what you can and cannot change
- Playwright MCP session blocked: four causes, four fixes
- Playwright MCP and captchas: what actually gets you past
- Cloudflare and a browser MCP server: what is being read
- Can an AI agent solve a captcha? The honest answer
- Getting an AI agent to fill out forms
- The best model for an MCP browser agent, and what it really costs
- Browser problem or model problem?
- How to use a Playwright MCP server with Claude Code
- Extracting data to a CSV with an AI agent
- Monitoring a page for changes with an AI agent
- How to let Claude Desktop control a browser
- How to add a browser to Cursor as an MCP server
- Using an AI agent to hunt for apartments
- Getting website data into Google Sheets with an AI agent
- Using an AI agent to download invoices from portals
- AI agents for web research
- Using an AI agent to test your own website
- How to add a browser to Cline as an MCP server
- Posting to social media with an AI agent
- Posting to Facebook with an AI agent
- Posting to Instagram with an AI agent
- Posting to X with an AI agent
- Automating LinkedIn posts: read this first
- Appointment bots: what they are and what an agent can legitimately do
- Track prices across sites with an AI agent
- Build a lead list with an AI browser agent
- Run an AI browser agent on a schedule
- AI browser agent with a local LLM: what changes
- Should you log your AI agent into your accounts?
- How to write a task an AI browser agent can follow
- Move data between two web apps with an AI agent
- The MCP server
- How the tools are shaped, and why
- Playwright MCP vs the Playwright CLI: which fits when
- Playwright MCP: browser is already in use, and the fix
- Playwright MCP best practices: four decisions that matter
- Playwright MCP with a proxy, and the three leaks it leaves
- A browser MCP server in GitHub Copilot: setup and limits
- Using a browser MCP server for web scraping: the pattern
- Which LLM for browser automation: the four properties
- How to build a browser agent, and what to take instead
- Getting an AI agent to log into a website: three routes
- MCP tools, resources and prompts: who controls each
- How many MCP tools is too many? The context arithmetic
- How to build an MCP server: the decisions, not the scaffold
- Local or remote MCP server: what changes, and what does not
- Writing an MCP client in Python: the thirty-line version
- Self-hosted AI agent: what one actually costs to run
- How long an AI browser agent takes per step, measured
- Text, HTML, snapshot or screenshot: what the agent should read
- Giving an AI browser agent a stopping condition
- Keeping an AI browser agent out of destructive actions
- Why did the AI agent click the wrong thing
- When the page changes under the AI agent
- Running one AI agent task across a list of sites
- Seeing a page as it appears in another country
- Getting data out of a dashboard with no export button
- Two browsers in one session: main and support
- Finding the dead links on a site with an AI agent
- Filling a CRM record from a company's website
- One form submission per spreadsheet row, with an AI agent
- Dated screenshots of a page as evidence
- Checking order and delivery status with an AI agent
- Reading a PDF that opens inside the browser
- Summarising a long page or thread with an AI agent
- Collecting every image on a page with its caption
- Collecting event and course listings with an AI agent
- Cancelling a subscription with an AI agent
- What an AI agent can and cannot do inside an iframe
- Shadow DOM and an AI agent: you can click it, you cannot read it
- How to add a browser to Codex as an MCP server
- What a page snapshot costs, per control
- How to let Gemini CLI use a browser
- Native selects and the ones that only look like selects
- Clicking by selector or by coordinates
- How long the agent waits before it gives up
- What a second browser costs
- Uploading a file with an AI agent, and why this one cannot
- Watching the agent work, and when it is worth it
- When not to use an AI browser agent
- Agent or script: deciding once instead of every time
- Using the keyboard instead of the mouse
- Secrets in an agent task: where they end up
- What an agent run should log
- Deduplicating what an AI agent collects
- Normalising values across sites
- Validating an AI agent's output
- Reading a table with an AI agent
- Driving a site's own search and filters
- The task works headed and fails headless
- How to let the AdaL CLI use a browser