[RFC] application/experiment intent type for controlled A/B tests (create, edit) #22
jacklyn-cheng
started this conversation in
RFC
Replies: 1 comment
|
Hi @adambarrus, whenever you have a moment, could you please take a look at this? |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
What type are you proposing?
application/experiment, with the standardcreateandeditactions — for apps that help merchants improve conversion through A/B testing.An experiment is an app-owned controlled trial that assigns storefront visitors to a control plus one or more variations by percentage and measures the outcome difference on a primary metric.
Why each action:
create— a merchant describes a test in conversation and Sidekick hands the merchant off to the app with a draft staged; they review and launch in the app. An A/B test changes what live visitors see, so creation must stay a handoff, never a command.edit— a test is a living object while it runs: merchants adjust the traffic split, rename variations, and pause or end the test. Lifecycle changes (pause, complete) ride oneditas astatustransition rather than becoming new actions.What do merchants ask Sidekick to do that needs it?
Real requests we see from merchants running A/B tests today are like:
With this intent type, each of those sentences becomes something Sidekick can act on.
For the first, Sidekick collects the missing details in conversation, then invokes
create— opening the app's test builder with the draft staged (variations, split, and end date prefilled) for the merchant to review and launch.For the other two, Sidekick finds the right experiment among the merchant's tests and invokes
editwith itsid— opening that experiment in the app with the new split, or thestatuschange, staged for the merchant to confirm. Without the type, none of these requests has an intent to resolve to.What apps would register intents for it?
Apps that help merchants split storefront traffic across experiments would all need this intent and the standard fields it defines.
A typical A/B testing app in the Shopify App Store offers tests across one or more of price, theme, shipping, content, and checkout — ABConvert (us), Intelligems, Shoplift, and Visually.io, among others. Each has a self-serve test-builder flow that
createmaps onto; and a running test is something merchants come back to, to pause it or end it, which is whateditmaps onto.Shopify's own Rollouts feature already uses the same control/treatment/traffic-percentage shape natively for theme and checkout changes.
What fields does it need?
Every controlled experiment shares one structure, whatever the variable under test: a set of variations (a control plus one or more challengers), a traffic split across them, and a lifecycle. That structure is universal across experiment types and across apps, so it is what the canonical schema should carry:
{ "type": "object", "properties": { "id": { "type": "string", "description": "The app's identifier for the experiment (used by edit)" }, "name": { "type": "string", "description": "Merchant-readable experiment name" }, "status": { "type": "string", "enum": ["draft", "active", "paused", "completed"], "description": "Lifecycle state. Transitions are an edit on this field." }, "variations": { "type": "array", "description": "The arms of the experiment, control included", "items": { "type": "object", "properties": { "name": { "type": "string", "description": "e.g. Control, Variant A, or a merchant rename like $49" }, "traffic_percentage": { "type": "integer", "description": "Share of traffic, 0-100; sums to 100 across variations" }, "control": { "type": "boolean", "description": "Marks the baseline variation" }, "description": { "type": "string", "description": "Free-text summary of the change this variation makes, in the merchant's words; the app confirms exact values in its UI" } }, "additionalProperties": true } } }, "additionalProperties": true }Two deliberate choices in this shape:
descriptionis free text. The change a variation makes (a price, a theme, a shipping rate) shares no parseable structure across experiment types, so each app interprets it.inputSchema, not in the shared schema.Anything close in the existing list?
Nothing in the current catalog models a randomized traffic split with a measured outcome.
The nearest neighbor is
application/campaign, which schedules a marketing send to an audience; it has no control group, no traffic split, and no measured comparison — the defining properties of an experiment.Thank you for considering this, happy to discuss more about this proposal!
All reactions