Replies: 21 comments
|
Adding a B2B perspective from Shopware, and building on the buyer-agent negotiation direction in #502. We are currently working on B2B agentic commerce and have implemented RFQ/quote workflows as a vendor extension. Our experience suggests two complementary layers:
For B2B, this is not an edge case. Pricing may depend on buyer identity and contracts, volume tiers, inventory allocation, delivery dates, payment terms, and human approval. These workflows commonly involve revisions, expiry, counteroffers, and asynchronous steps before an order can exist. We would support a minimal, optional term-formation artifact or capability that works with bilateral negotiation as well as RFQ/procurement flows, without trying to standardize negotiation strategy or supplier selection. It should be revisioned and time-bound; bind parties, line items, and commercial terms; record acceptance or award state; and bind deterministically to the authoritative UCP Cart or Checkout. An opaque reference may help bridge existing implementations, but portable semantics should remain the destination. We would be glad to share concrete B2B flows and lessons from our extension if that helps shape a proposal or working group. |
|
Thanks Juan — the Shopware implementation experience is particularly useful here because it gives us a concrete B2B case rather than designing this boundary from first principles. I also like the separation you describe between the two layers. #502 can stay focused on whether and how a Business advertises and operates negotiation, while this thread can focus on what must survive once commercial terms have actually been agreed and need to enter authoritative UCP transaction state. Rather than choosing a schema yet, I think the next useful step is to derive the smallest common contract from a few concrete flows. If you can share sanitized examples from your extension, especially an RFQ → revised quote → accepted terms → Cart/Checkout flow, an expiry/revision or counteroffer across an asynchronous pause, and a flow involving human approval, I can turn those into protocol-level test vectors and compare the two handoff models discussed here: an opaque reference versus a minimal portable artifact. The comparison I’d want to make is whether each model can preserve term identity/revision, parties, item/quantity scope, expiry, acceptance or award state, and deterministic binding to the Business-authoritative Cart or Checkout without pulling negotiation strategy or supplier-selection semantics into UCP. If those flows expose a stable common subset, that would give us a much stronger basis for a vendor-namespaced experiment or eventual proposal than designing the artifact top-down. |
|
Hi @arjun2075 Sharing more info below. You can see an example of our current B2B Quote capability here:
It currently supports request, retrieval/polling, counter, accept, and decline operations. The lifecycle distinguishes the merchant turn ( It takes into account the three flows you named:
One implementation lesson already seems relevant to the handoff boundary: the quote is a durable, customer-authorized commercial object, not just a Cart with a temporary discount. Its identity, ownership, state, expiry, offered unit prices, tax-status-aware totals, and decision trail have to survive that human-led pause. In our current implementation, acceptance places an order through the B2B quote flow; a UCP-native handoff would instead need to bind that accepted commercial state deterministically to the resulting Cart or Checkout. Let me know if this helps |
|
This helps a lot — thanks Juan. The fact that the quote has to survive async and human-led pauses as its own durable commercial object is particularly useful for defining the boundary here. I’m going to work through the spec and the three flows you shared and map the accepted-quote → Cart/Checkout handoff against the two models we discussed: carrying an opaque reference versus carrying a minimal portable set of agreed-term semantics. I think the useful test is whether we can preserve quote identity/revision, expiry, item/quantity and price scope, acceptance state, and deterministic binding into the resulting transaction without pulling the negotiation workflow itself into Checkout. I’ll bring the resulting comparison back here rather than proposing a schema upfront. |
|
One property worth adding to that comparison, because it decides whether either model is safe rather than whether it round-trips. Everything on the list so far is preservation: term identity and revision, parties, item and quantity scope, expiry, acceptance state, deterministic binding. All necessary. But a handoff can preserve every one of those and still move pricing authority to whoever formed the terms, because preservation describes what the artifact carries, not what the Business is permitted to do with it. The missing property is direction of authority. When the accepted quote reaches the authoritative Cart or Checkout, is it an input the Business reads a price out of, or an assertion the Business must independently re-derive and agree with before the transaction exists? Those two can produce an identical wire format and have very different failure modes. In the first, anyone who can mint or replay a term-formation artifact can set a price. In the second, a forged or stale artifact fails closed because the Business recomputed and disagreed. Juan's point that the quote is a durable, customer-authorized commercial object is exactly why this bites. Durable objects outlive the conditions they were priced under. Inventory allocation moves, a contract tier lapses, a delivery date slips past what the quoted freight assumed. Expiry covers the clock but not the state, so an unexpired quote can still be one the Business would no longer agree to, and the human-led pause is precisely the window where that drift happens. Concrete test for both models: can the Business decline an artifact that is valid, unexpired, correctly bound to the right parties and line items, and still wrong? If a model has no representation for "I recognize this and I do not agree", it has quietly made the external artifact authoritative, and the opaque-reference version hides that more effectively than the portable one, since there is nothing to re-derive against. That also suggests a fourth flow worth adding to the three you asked Juan for: accepted quote, then a Business-side rejection at binding time. It is the flow that separates the two models, and it is the one neither an RFQ nor a counteroffer example exercises. |
|
That’s a useful distinction. I was treating deterministic binding mainly as an identity/integrity property, which misses who actually has authority over the terms at the transaction boundary. I’ll add the fourth flow: an accepted, unexpired, correctly scoped quote reaches Cart/Checkout after relevant Business state has changed. One thing I’d like the comparison to distinguish, though, is whether the artifact represents a Business-issued commitment or merely accepted terms that still require Business revalidation. In the latter case, “recognized but no longer agreed” clearly needs to be representable. In the former, allowing arbitrary repricing at binding time could undermine the semantics of the quote itself. So I think the additional property is not just rejection support, but explicit authority/commitment semantics: who authorized the terms, what authority the artifact carries, and under what conditions that authority can be invalidated before execution. I’ll make that a separate dimension in the comparison rather than folding it into deterministic binding. |
|
I worked through the implementation cases we discussed and turned them into a small review artifact with seven protocol-level vectors: https://github.com/arjun2075/ucp-term-handoff The vectors cover accepted RFQ, async revision, human approval, scope mutation, expiry, Business rejection at binding, and a Business-issued firm commitment. The main thing that fell out of the exercise was that four concerns need to stay separate:
The V6/V7 pair was particularly useful. V6 models a valid, recognized artifact that may still be rejected because its commitment semantics explicitly retain Business revalidation. V7 models the opposite case: under the harness's proposed Comparing the two handoff models also exposed different gaps:
A hybrid — portable semantics plus an authoritative Business reference and classified binding result — looks worth testing, but I don't think the evidence is strong enough yet to make that a UCP proposal. @westonale @juanfernandez — I'd particularly like to know whether the commitment distinction and the rejection taxonomy match the failure/implementation cases you had in mind, or whether there's a case these vectors still miss. |
|
Both match, and I7 through I11 are sharper than what I described. Separating Two cases the vectors do not reach. Both come from one assumption, visible in I12: binding is one evaluation, at one instant, of one artifact, against one target. The window does not close at binding. In both V6 and V7 the state change lands before the bind attempt, and the timeline ends at the bind. But a bound Checkout is not an executed transaction. It sits until completion, and the same conditions can move again in that gap: V6's inventory allocation, V7's contract tier, or a freight assumption the quote priced. By then the commitment has been evaluated and there is nothing left to revalidate against, so the harness has no outcome for "bound, then the conditions that made it bindable stopped holding". Two clocks are in play and no invariant says which governs the interval between them: the quote's expiry and the Checkout's own lifetime. It bites hardest on The shape that sidesteps that interval rather than shortening it is worth naming: where settlement is conditional on fulfillment, the terms lock at agreement and value does not move until delivery is confirmed, so the interval does not have to be drift free. That converts a revalidation problem into a release condition. Not the only answer, but it is the one that makes the gap between bind and execute safe rather than short. Binding is all or nothing. The taxonomy in unresolved question 3 is per artifact. I3 forbids silent scope expansion and nothing covers scope contraction, so a 40 line award where 38 lines still bind and 2 are no longer available resolves either as a whole artifact |
|
Thanks — I extended the artifact around both gaps you identified, while leaving V1–V7 unchanged: https://github.com/arjun2075/ucp-term-handoff For the bind → execute gap, I added two cases:
For partial binding, I ended up going slightly broader than per-line results. A line isn't always independently severable when pricing, freight, minimums, bundles, or thresholds span multiple lines, so the harness now uses explicit binding units / atomicity groups:
That also pushed idempotency down to the granularity of artifact + revision + binding unit + target transaction. The model comparison still doesn't select a universal carrier. Both opaque references and portable artifacts need additional transaction-lifecycle semantics for the post-bind cases, and partial binding needs an authoritative grouped-result contract either way. The parts I'd most like you to challenge are:
|
|
Binding units are the better abstraction. Per-line results are just the case where every unit has one line. Taking your three questions in order, against what the harness does today: 1. The four modes are distinct, but they sit on two axes, not one. Three of them say what bounds execution: the original term expiry, a separate deadline, or the transaction's own lifetime once authority transfers at bind. Conditional release answers a different question, which is what has to happen before value moves. Real terms often need both, because a delivery-gated release still needs a date it cannot pend past. A single 2. Units capture it, with one coupling V10 avoids by construction. None of V10's 40 lines is an order-level term, and its commercial terms state outright that the four units are independent. In a real award, freight, an order minimum or a volume tier is often computed across the whole award, so when U4 fails the price basis for U1 through U3 can move with it. With 3. Yes under the same revision, never under the same attempt. The idempotency key identifies an attempt and the revision identifies an authorization, and the harness currently ties the two together. Once U4 has a prior failure it stays rejected even after its lines come back, which is exactly what |
|
Weston, your validation of binding units helped settle that abstraction. I’ve opened a reviewable revision: arjun2075/ucp-term-handoff#1 Deadline policy now composes with optional release conditions, including an explicit pending-at-deadline disposition. Cross-unit contraction uses previously authorized deterministic adjustment rules. Attempt identity is separate from revision: same-attempt replay preserves the failure, while a fresh attempt can recover transient availability under the same valid authorization. Could you challenge those three choices—particularly the deadline boundary/disposition, the limits on authorized contraction adjustment, and the transient-versus-structural retry classification? V1–V7 remain unchanged. |
|
I'm coming from the world of hotels, and would like to suggest adding a third flow that differs from the B2B cases above in one respect: the terms are formed by a party that is neither the buyer nor the Business. PR #780 (Lodging Booking Capability) does this as an existence proof of need. The spec separates "provisional discovery" from "authoritative booking": search, quotation and room lookup return "provisional rates", available room types and policy summaries,. Creating the booking session is where the Business locks or evaluates real-time inventory, resolves binding rate rules, calculates totals and attaches authoritative cancellation terms. What between these stages is minimal. The Platform sends In the taxonomy above that sits close to model 1, with a trace of model 2: ordinary Checkout inputs plus identifiers whose provenance is out of band, with the Business restating the terms. It is cheap and it works. What it does not preserve maps onto the list you and Juan assembled:
In the world of lodging this is a known, prevalent, annoying problem: a website will show a low rate to get the click and then show the user a higher price when they're about to book. Having run the Price Accuracy team while part of Google Hotels, I've been fighting this problem for years. I'd like to prevent it in the world of Agents. The reason I think this bears on the general boundary rather than only on lodging: it is an existing capability where external term formation is already assumed and deliberately left unspecified. If a minimal portable artifact lands, lodging is a test of whether it can describe terms formed by a third party rather than by the two parties to the transaction. If it cannot, model 3 ends up meaning bilateral only. |
|
This is a useful third case, and I think it exposes a boundary slightly earlier than the accepted-term handoff we’ve been modeling. In the lodging flow, room_rate.id or room_type.id + rate_plan.id + itinerary tells the Business what to evaluate, but it does not identify the actual provisional price/policy the buyer saw. The Business then authoritatively validates inventory and evaluates/restates the commercial terms when the booking session is created. So I think there are two separate things we should avoid conflating:
The issuer / authorized_by distinction can represent different provenance and authority parties. The gap in the current harness is different: it assumes the artifact presented at the binding boundary is an accepted/awarded, authority-bearing artifact. Your lodging example suggests another semantic class—an authentic, attributable observation that is intentionally non-binding. The question I’d test from your example is: should the handoff carry the actual provisional observation the buyer saw—price/policy plus enough source/time identity to compare it with the Business-authoritative session—or is the intent that only the room/rate identifiers cross that boundary? If it is the former, that seems like a distinct provenance case rather than another Business-authorized quote case, and it would let us represent the price-accuracy delta without implying that the discovery source had authority over the Business. |
|
should the handoff carry the actual provisional observation the buyer
saw—price/policy plus enough source/time identity to compare it with the
Business-authoritative session
or
only the room/rate identifiers cross that boundary
The former is what I would recommend. My rationale here is that providing
only the room/rate identifiers leaves open the bait-and-switch dynamics
we've seen to date. They may be less annoying to Agents than to humans,
but if we have an opportunity to design them out, I believe we should
attempt to do so.
…On Sun, Sep 13, 2026 at 9:10 PM Arjun ***@***.***> wrote:
This is a useful third case, and I think it exposes a boundary slightly
earlier than the accepted-term handoff we’ve been modeling.
In the lodging flow, room_rate.id or room_type.id + rate_plan.id +
itinerary tells the Business what to evaluate, but it does not identify the
actual provisional price/policy the buyer saw. The Business then
authoritatively validates inventory and restates/recomputes the commercial
terms when the booking session is created.
So I think there are two separate things we should avoid conflating:
1. provenance of a provisional commercial observation — what terms
were shown, when, and by/source from whom; and
2. authority to bind those terms — which the provisional observation
may not have at all.
Our current issuer / authorized_by split helps with the second question,
but the harness is still built around artifacts presented for binding. A
third-party lodging quote may instead be authentic and attributable while
remaining entirely non-binding.
The question I’d test from your example is: should the handoff carry the
actual provisional observation the buyer saw—price/policy plus enough
source/time identity to compare it with the Business-authoritative
session—or is the intent that only the room/rate identifiers cross that
boundary?
If it is the former, that seems like a distinct provenance case rather
than another Business-authorized quote case, and it would also give us a
way to represent the price-accuracy delta without implying that the
discovery party had authority over the Business.
—
Reply to this email directly, view it on GitHub
<#812?email_source=notifications&email_token=BKCATR6PFRD2Y7KJNE6UQAL5O5AQ5A5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBUGI3TGMRSUZZGKYLTN5XKOY3PNVWWK3TUUVSXMZLOOSWGM33PORSXEX3DNRUWG2Y#discussioncomment-18427322>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/BKCATR7N4KRZYEHCNWY5PSD5O5AQ5AVCNFSNUABJKJSXA33TNF2G64TZHMYTCMRVGU4TCOBUGE5UI2LTMN2XG43JN5XDWMJQG43TCNRYGOQXMAQ>
.
Triage notifications, keep track of coding agent tasks and review pull
requests on the go with GitHub Mobile for iOS
<https://github.com/notifications/mobile/ios/BKCATRZL7IGKTIDC64U6AYT5O5AQ5A5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBUGI3TGMRSUZZGKYLTN5XKOY3PNVWWK3TUUVSXMZLOOSVGM33PORSXEX3JN5ZQ>
and Android
<https://github.com/notifications/mobile/android/BKCATR5TWQGJHRRYM6AMZYT5O5AQ5A5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBUGI3TGMRSUZZGKYLTN5XKOY3PNVWWK3TUUVSXMZLOOSXGM33PORSXEX3BNZSHE33JMQ>.
Download it today!
You are receiving this because you commented.Message ID:
<Universal-Commerce-Protocol/ucp/repo-discussions/812/comments/18427322@
github.com>
|
|
That clarifies it. I agree the useful object here is the provisional commercial observation itself, not just the lodging identifiers. I’d keep its semantics deliberately weaker than an accepted-term artifact: it can preserve what was shown, its source/provenance, observation time, commercial scope, price/policy terms, and any stated validity — but it carries no implied authority over the Business. Any buyer acceptance or authorization would need to be explicit and would not by itself turn the observation into a Business commitment. Then the Business-authoritative booking session can be compared against that observation explicitly. A price/policy change becomes an attributable delta rather than the earlier terms simply disappearing at session creation. One distinction I’d keep separate is observation from enforcement: preserving the provisional terms makes bait-and-switch detectable, but the protocol would still need an explicit policy if a Platform is expected to block, warn, re-confirm, or otherwise react to a material delta. I’ll model this as a separate provenance vector after the current accepted-term PR review, rather than folding it into the binding-authority cases. |
|
Read PR #1 at 1. The exclusive boundary is right. Late observation needs a finality point. Nothing bounds how late an earlier verified event can arrive, so the disposition applied at the deadline is never final. V13 with 2. The authorization limits hold. The repricing guard compares against the wrong price. Attempt history records that a unit bound, not the price it bound at, so the check compares the tier price with the artifact's list price. That misfires in both directions. Replay V12 against the four history entries its own result implies (U1 to U3 3. The classes are right for the unit reasons, but the class should follow who has to act. A per-unit |
|
Weston — I pushed the follow-up to PR #1 at I addressed the three cases from your review:
I also changed the initial deadline tests so they clear persisted terminal state and exercise the deadline calculation itself. The PR now reports 70/70 tests with V1–V7 still byte-for-byte unchanged. Could you take another look at those three closures, particularly whether the explicit adjustment representation preserves the distinction you had in mind between existing authorization and a separate economic adjustment? |
|
Read e7dd449c. The adjustment has the right shape: bind-time prices stay in history, the accepted tier produces a separate record with the previous and new basis, and a change the table does not cover fails instead of becoming a new effective price. That distinction holds for the request that emits the adjustment. It does not survive the next evaluation, because nothing records that the adjustment was applied. Take the restock request from It also fails a later bind that reprices nothing. Add a 2-unit row at 1100, and bind the units one at a time by presenting all four each request with the rest marked unavailable. When U2 binds, count reaches 2 and L01 is adjusted 1200 to 1100. When U3 binds, count is 3, the tier is still 1100, so nothing new is repriced, yet the request fails: L01 is still recorded at 1200 while the authorized price for two prior units is now 1100. In both cases the comparison reads the bind-time price where it needs the current basis, the bind-time price plus the adjustments already applied. Recording each applied adjustment as its own history entry closes both, as long as the entry carries what caused it: the triggering attempt id, the tier row applied, and the revision. The evaluator then compares the current basis with the prior authorized price, a replay recognizes an adjustment it already applied instead of rejecting or re-emitting it, and a consumer can deduplicate. History is never rewritten, and after the fact an adjustment stays distinguishable from a new authorized transition, which is the distinction you asked about. Two smaller points:
|
|
Thanks — I pushed a follow-up in Accepted commercial-basis adjustments are now persisted as append-only history. Bind-time prices remain unchanged; subsequent evaluation derives the current basis from the bind-time price plus the applied adjustment chain. That fixes both adjustment cases you described:
Each adjustment records the artifact/line, previous and new basis, triggering attempt, tier row, and revision. Replay uses one canonical identity for those fields, so an already-applied adjustment is not emitted or persisted again. The basis derivation also validates each link against the running basis; inconsistent or reordered adjustment history fails closed rather than producing an effective price. I also addressed the two smaller points:
Validation is now 82/82 tests, 14/14 vectors, unchanged V1–V7 hash guards, and clean |
|
Read 31e11d59. Both adjustment traces now pass, and deriving the current basis from bind-time prices plus the applied chain is the right model. The terminal outcome check accepts on the label rather than the outcome. The conflict check runs only when the record's reason is one of the two deadline dispositions and a release condition exists, and it then compares only the reason, so three records replay as executed:
invariants.md says the persisted record is validated against the authorized disposition, and it is, by its reason. What is not checked is that Smaller: a broken or reordered adjustment chain returns |
|
Thanks — your three terminal-outcome cases were valid, and the broken-chain recovery point was valid as well. I ended up tightening both areas beyond the individual reproductions rather than special-casing the vectors. The current implementation is in Terminal outcome replayThe persisted terminal record is no longer accepted on its Replay now validates the complete semantic outcome:
against the terminal class permitted by the accepted policy and the record's timing. That closes the cases you identified:
For release finality, I added a minimal That means a legitimate terminal result is final across later attempt-state drift: disappearing or contradictory fields in a later attempt do not retroactively re-decide the historical result. Transient evaluator results are also explicitly outside terminal history. Things such as There is one trust boundary worth stating explicitly. The terminal record and its release basis are trusted Business-store history; the basis is not separately authenticated from the record carrying it. Replay detects internal, policy and timing contradictions, but it cannot detect coherent replacement of the entire trusted record plus its provenance. Also, Commercial-basis historyI also changed the broken-chain behavior you called out.
If the history is available but internally inconsistent, it is structural invalidity and returns The current basis is still derived from immutable bind-time prices plus the accepted append-only adjustment chain; bind-time prices themselves are not rewritten. The history validation is now independent of whether the current request happens to replay as For each artifact revision, persisted adjustments are validated in recorded order against authoritative successful-bound history. Among other things, the validator now checks:
Successful bound records used as authoritative history are also validated before they can contribute basis/count/trigger state, including an exact match between Adjustment replay keeps the idempotency key separate from the payload. An already-applied identity suppresses re-emission only when the persisted effect and transition provenance match the transition being replayed. For later adjustments in a chain, that comparison uses the basis immediately preceding that record — e.g. There is a narrower remaining trust boundary on historical adjustment triggers: attempt history does not retain ordering, so among otherwise structurally compatible successful historical attempts the harness cannot independently prove which one occurred at the exact transition count. Membership, lines, target, successful status, count, tier, price, source and cross-record context are checked; historical ordering itself is not reconstructed. ValidationCurrent validation is:
No vector or golden value changed. Thanks for pushing on the distinction between a matching disposition label and an actually authorized persisted outcome. That ended up exposing the broader replay/history issue, and the implementation is now checking the structured historical state rather than relying on labels or the current request to reconstruct it. |
Uh oh!
There was an error while loading. Please reload this page.
Discussion #502 explores buyer-agent negotiation. This question is slightly broader: regardless of whether terms come from bilateral negotiation, RFQ, procurement, or an external market mechanism, what artifact—if any—should cross into UCP before Checkout?
UCP currently appears to provide three relevant pieces:
The boundary I am trying to isolate is different from #738/#773. It concerns commercial terms that may be formed outside UCP—or before a UCP Cart or Checkout exists—and then need to be carried into a transaction with the Business responsible for executing them.
Examples include:
Conceptually:
external term formation
↓
terms selected or agreed
↓
handoff into UCP
↓
Business-returned Cart or Checkout
↓
Order
I do not think UCP should define bidding rules, negotiation strategy, participant selection, auction clearing, scoring or ranking, reserve prices, or how a marketplace reaches a result. Those concerns can remain mechanism-specific.
The interoperability question is the handoff: should different term-formation mechanisms be able to produce something that a UCP Platform and Business can consistently recognize, validate, and bind to the resulting transaction?
I see three plausible boundaries:
Questions for maintainers:
This is not proposing a competing negotiation design or an auction protocol. #502 should remain the discussion for the bilateral negotiation flow. I am trying to isolate the more general architectural boundary before proposing any schema or implementation.
If maintainers consider this the same scope as #502, I am happy to consolidate the discussion there.
All reactions