This repository was archived by the owner on Feb 3, 2026. It is now read-only.
Releases: zanminwang/webtask
Releases · zanminwang/webtask
Release list
v0.27.0
Changes
-
Standardize time parameters to seconds: All time-related parameters now use seconds consistently
typing_delay: 80ms → 0.05swait_for_load/wait_for_network_idletimeout: 10000ms → 10.0skeyboard_typedelay: 80ms → 0.05s
-
Add
raise_on_timeoutparameter:Agent.wait_for_load()andAgent.wait_for_network_idle()now acceptraise_on_timeout(default: True)- When False, silently returns after timeout instead of raising TimeoutError
- Useful for lenient waits where you want to continue regardless
Example
# All time params now in seconds
agent = await wt.create_agent(
llm=llm,
wait_after_action=1.0, # seconds
typing_delay=0.05, # seconds (was 80ms)
)
# Lenient wait - continue even if timeout
await agent.wait_for_network_idle(timeout=5.0, raise_on_timeout=False)v0.26.0
Changes
- Live form value parsing: DOM context now captures live input values (
inputValue,textValue), checkbox/radio states (inputChecked), and select option states (optionSelected) from CDP snapshots - Configurable typing delay: Added
typing_delayparameter to Agent (default: 80ms) - Centralized constants: Added
constants.pywithDEFAULT_WAIT_AFTER_ACTIONandDEFAULT_TYPING_DELAY
Bug Fixes
- Fixed issue where user-typed values (e.g., search queries) weren't visible in DOM context
- Fixed 17 deselected tests by adding missing
@pytest.mark.unitmarkers
v0.25.1
Bug Fixes
- Fix TypeTool clear parameter not working: The
clearparameter in thetypetool was usingelement.fill("")which would lose focus beforekeyboard_type. Now passescleartokeyboard_typewhich uses JavaScript to clear the active element while maintaining focus.
Other Changes
- Add demo video example and GIF to README
- Add
videos/to gitignore
v0.24.0
Changes
- TypeTool and TypeAtTool: New tools that click-then-type in one action for more human-like interaction
- SelectTool: New tool for dropdown selection
- Simplified message types: Single Message class with Role enum instead of separate UserMessage/AssistantMessage classes
- Tools receive wait_after_action directly: Cleaner tool initialization without browser.wait() calls
- ToolParams base class: Rejects extra parameters for stricter validation
- Fixed clear functionality: Improved action timeout handling
Breaking Changes
- Removed test recording/replay infrastructure (
webtask.testingmodule) - Removed e2e test workflow from CI
v0.23.2
Changes
- Fast-fail timeout for element actions: Element actions (click, fill, type, upload) now use a 100ms timeout instead of Playwright's 30s default. This prevents long waits when clicking non-existent or hidden elements while still performing actionability checks.
Details
- Added
DEFAULT_ACTION_TIMEOUT = 100msinplaywright_element.py - Applied timeout to
click,fill,type,upload_filemethods - Removed context-level
set_default_timeoutfromplaywright_browser.py - Deleted unused
constants.py
v0.23.1
Changes
- Add
ToolParamsbase class withextra="forbid"to reject invalid LLM tool parameters - All tool Params classes now inherit from
ToolParams - Export
ToolandToolParamsfromwebtask.llm - Add test for extra field rejection
- Move TODO from README to
docs/todo.md
v0.23.0
Changes
Human-like interaction
- Remove
FillToolandTypeTool(element-level typing) - Remove
TypeTextAtToolfrom pixel tools - Add
KeyboardTypeTool(page-level keyboard typing) - User flow: click element to focus, then type on keyboard
Improved scrolling
ScrollDocumentToolnow scrolls 50% of viewport (maintains context, won't cut elements in half)- Uses JavaScript
window.scrollBy()instead of PageUp/PageDown
Tool organization
- DOM mode: click, upload + keyboard tools
- Pixel mode: click_at, hover_at, scroll_at, scroll_document, drag_and_drop + keyboard tools
- Keyboard tools (both modes): type, key_combination
v0.22.1
Changes
- Keep
li,ul,olelements as semantic HTML tags - Dropdown menu items now get element IDs (e.g.,
li-0,li-1) - Add tests for semantic knowledge functions
v0.22.0
Changes
- Simplify agent modes from three (text/visual/full) to two (dom/pixel)
- dom mode (default): Uses element IDs for interaction, includes DOM + screenshots
- pixel mode: Uses screen coordinates, includes screenshots only
- Move ScrollDocumentTool to pixel mode only
- Both modes now get screenshots for better LLM context
v0.21.4
What's New
- Added
[END OF PAGE]marker: DOM context now ends with a clear marker to help the LLM understand where page content ends, preventing it from echoing back DOM elements in responses.