Skip to content

stagehand-python@4.2.0a0.dev1541

@seanmcguire12 seanmcguire12 tagged this 26 Sep 19:49
thanks @mikhail-koviazin for the contribution here!

## why

The composed-tree XPath parser evaluates `text()` and `.` through the
same helper, `element.textContent`. In XPath these are not the same
thing: `.` is the string-value of the element, so it covers the whole
subtree, while `text()` is the node-set of the element's **direct child
text nodes**.

This parser is not a rare path. It takes over whenever the document
contains a shadow root anywhere, so a single unrelated web component on
the page changes what a locator matches, silently and with no error.

Measured against `document.evaluate()` on a build of `main`
(`a73da68b`), fixture served over http. The only difference between a
run that matches native and a run that does not is one unrelated
`attachShadow()` call elsewhere in the same document:

| XPath | `document.evaluate()` | Stagehand |
| --- | --- | --- |
| `//button[text()='Save']` | 1 | 2 |
| `//div[text()='a']` | 1 | 0 |
| `//div[contains(text(),'b')]` | 0 | 1 |
| `//div[@id='split'][text()='y']` | 1 | 0 |
| `//div[@id='split'][contains(text(),'y')]` | 0 | 1 | |
`//div[@id='split'][normalize-space(text())='x']` | 1 | 0 | |
`//button[.='Save']` | 2 | 2 (control: `.` is already correct) |

Fixture: `<button id="wrapped"><span>Save</span></button>` before
`<button id="direct">Save</button>`, plus `<div
id="mixed">a<span>b</span></div>` and `<div id="split">x<br/>y</div>`.

The count is not the worst part. The button whose label sits inside a
`<span>` comes first in document order, so `//button[text()='Save']`
returns it first and `.first().click()` clicks the wrong button. Nothing
throws, nothing logs, and the run continues on the wrong element. The
other direction is a silent zero on markup as ordinary as
`a<span>b</span>`.

## what changed

`text()` now reads its own node-set, and `.` keeps exactly the meaning
it has today:

- `text()='v'` is true when **any** direct child text node equals `v`.
Comparing a node-set to a string is existential in XPath.
- `contains(text(),'v')` and `normalize-space(text())='v'` read the
string-value of the **first** node, because both functions take a
string, so the node-set collapses.
- `.='v'`, `contains(.,'v')` and `normalize-space(.)='v'` keep reading
the string-value of the element, unchanged.

Mechanically, in `packages/extension/dom/locatorScripts/xpathParser.ts`:
the three patterns that accepted `(?:text\(\)|\.)` now capture which
token they matched, and `textEquals` / `textContains` carry a `source:
"self" | "text"` field that `evaluatePredicate` reads. The field is
optional and absent means `self`, so predicates constructed anywhere
else keep their current behavior. `and`, `or` and `not` carry it through
without changes.

## test plan

`packages/extension/tests/xpath-text-predicates.test.ts` (new, unit):
the parser keeps `text()` and `.` apart for `=`, `contains()` and
`normalize-space()`, and carries the source through `or` and `not`.

`packages/sdk-ts/tests/integration/locatorXPathTextPredicates.test.ts`
(new, integration): the table above against real Chrome. It has to be
served over http, since `data:` URLs stay on the native engine and never
reach this parser at all. Registered in `scripts/test-integration.ts`
next to the other locator groups.

Both are red on `main` and green with the change. On `main` the
wrong-button case fails with `expected 2 to be 1` and the silent-zero
case with `expected +0 to be 1`.

Context: #1679 consolidated the parser copies and #1683 taught this one
`text()`, `contains()` and `normalize-space()`. This is the same layer,
one step further in.


<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Fixes the composed-tree XPath parser so `text()` reads only direct child
text nodes. Previously it read `element.textContent` (same as `.`),
causing diverging matches and wrong clicks on pages with any shadow
root.

- Old: `text()` behaved like `.`. New: `text()` evaluates direct child
text nodes; `.` still reads the element's subtree string-value. This
aligns with `document.evaluate()`.
- If a locator relied on `text()` to match nested text, use `.` instead.
- `text()='v'` is existential across child text nodes;
`contains(text(),'v')` and `normalize-space(text())='v'` collapse to the
first text node.
- Adds unit and integration tests; integration runs over http and is
registered in `scripts/test-integration.ts`.

<sup>Written for commit 482e0a947b59752e07b74efc8110edac2346858a.
Summary will update on new commits.</sup>

<a
href="https://cubic.dev/pr/browserbase/stagehand/pull/3052?utm_source=github"
target="_blank" rel="noopener noreferrer"
data-no-image-dialog="true"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source
media="(prefers-color-scheme: light)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img
alt="Review in cubic"
src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a>

<!-- End of auto-generated description by cubic. -->

---------

Co-authored-by: Mikhail Koviazin <mikhail.koviazin@gmail.com>
Assets 2
Loading