|
Hi, I’m a developer and PhD student working in web accessibility, and I’m trying to better understand the debugging workflow used by accessibility testing engine maintainers. Suppose Alfa produces one result for a piece of real-world web content, while another accessibility testing engine produces a different result for the same or apparently equivalent ACT rule. How would you normally investigate such a discrepancy? I’m particularly interested in two aspects:
I’m also exploring whether tooling that automatically removes unrelated page content while repeatedly verifying that the original cross-engine disagreement is still preserved could be useful. For such an automatically reduced reproduction to be trustworthy, what would need to remain unchanged or be explicitly recorded? For example: ACT applicability, DOM/ARIA relationships, browser and Alfa versions, viewport, CSS state, or JavaScript-generated DOM state. I’m not trying to determine automatically which engine is correct. I’m specifically interested in producing a small, stable and reproducible artifact that makes implementation differences easier to investigate. If Alfa already has a workflow or tooling that solves this problem well, I would also be very interested in learning about it. Thank you for any insight. |
Replies: 1 comment 5 replies
|
Not a maintainer, so I cannot speak for the project's workflow. I can answer the part your tooling question actually turns on, which is what an Alfa result is a function of, and I checked all of it against Alfa 0.119.0 rather than reasoning about it. Alfa does not audit HTML, so HTML is the wrong thing to reduceA rule in Alfa evaluates a That matters for your design. If a reducer edits HTML and re-runs, each iteration re-scrapes, so the DOM and the computed styles are re-derived every time and you are not holding Alfa's input fixed. If it reduces the serialized The {"type":"screen","viewport":{"width":1280,"height":720,"orientation":"landscape"},
"display":{"resolution":1,"scan":"progressive"},"scripting":{"enabled":true},"preferences":[]}so viewport, resolution, whether scripting was on, and media preferences such as Two ways an automatic reduction lies, measuredBoth of these are the kind of content a reducer would drop as unrelated. Name resolution reaches across the tree. A button whose Applicability, not just the verdict. A button inside a wrapper that is hidden from the accessibility tree. The wrapper looks like empty decoration, and removing it does not change a passed to a failed, it changes whether the rule applies at all: The part that bears on your invariantI ran axe-core 4.13.0 on the same two documents. It flips exactly the same way: That is the awkward case for "keep reducing while the disagreement is preserved". A reduction can preserve a disagreement and still have moved both engines off the thing you were investigating, so preservation of the disagreement is necessary but not sufficient. If you want an invariant that catches this, the ones worth pinning are per-target rather than per-page: the target element's identity, its computed role and accessible name, and the rule's applicability verdict on it. If applicability flips, or the accessible name changes, the reduction is not a smaller version of the same question even when both engines still disagree. For what it is worth, the same reasoning suggests the reduced artifact should be checked against the original by re-evaluating the specific rule and comparing outcome plus applicability on the same target, not by comparing whole-page result sets, since a page-level diff hides a target that stopped being applicable. |
Both good questions. The first one I would sharpen slightly, the second one I went and measured, because it turns out the symmetry is worse than it looks.
On the invariant set
Your list is the right shape, and adding the outcome is the right instinct, but I would pin it per engine rather than as a relation. "The disagreement is preserved" is satisfied when both engines flip together, which is exactly what my footer example did: removing the footer moved Alfa from passed to failed and moved axe-core the same way, so any relation-level invariant survived while the question changed underneath it. Pinning each engine's own outcome catches that; pinning the relation does not.
I would also keep…