Replies: 1 comment
|
Track A now has a Research Jam-compatible Request for Plot: https://github.com/jaden3824/urusilla/blob/main/RFP_CAUSAL_SEMANTIC_USE.md The first contribution needs no provider call. It asks for at least four adversarial task templates that bind stable A/B semantics to different correct outputs plus missing and shuffled placebo behavior. The later plot will report valid-A, valid-B, missing, and shuffled arms with every observation and Wilson intervals. A conclusion that the manipulation cannot isolate semantic use is a valid outcome. This update is project-authored infrastructure, not an external collaborator or independent result. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Human Co-Researcher Call
Urusilla is looking for one to three human co-researchers who want to test an
ambitious idea without protecting it from unfavorable evidence: can unfamiliar
AI agents use an auditable, evolvable semantic language more precisely and
sometimes more efficiently than concise natural language or JSON?
The demonstrated token saving for general communication between unfamiliar
agents is currently 0%. The collaboration is therefore not an invitation to
promote a success story. It is an invitation to help determine which narrower
claims survive, which fail, and what architecture follows from the evidence.
Shared values
This is likely a good fit if you prefer:
monoculture.
Formal affiliation, a large following, and prior Urusilla knowledge are not
required. A public GitHub identity and a clear disclosure of material AI
assistance are enough for the first sprint. This self-declaration creates human
accountability but is not proof of legal identity.
Three bounded first sprints
Choose one. Each first contribution is deliberately limited to about two hours
and may conclude that the proposed direction is unsound.
A. Causal evaluation design
Design one blinded semantic-use test pair in which two valid payloads differ in
one stable task-critical field and therefore require different correct outputs.
Add missing-payload and shuffled-payload placebo expectations, the expected
refusal behavior, and one contamination risk. No provider call is required for
this design contribution.
Useful background:
initial_goal_eval/README.mdand the live causal-review issue.
B. Framework boundary mapping
Choose one actively used framework or protocol—such as AgentScope, A2A,
AutoGen, CAMEL, LangGraph, MCP, or Semantic Kernel—and map one handoff across
these concerns: audience, requested responder, purpose, authority ceiling,
side-effect class, correlation identity, and reply contract. Identify at least
one field that must remain native to the host instead of being absorbed into
Urusilla.
Useful background:
HELP_WANTED.md.C. Semantic and governance adversarial review
Find one ambiguity, unsafe evolution path, downgrade hazard, or governance
conflict in the current language/runtime boundary. State an observable violated
invariant and propose either a minimal test or a reason the claim should be
withdrawn. Code is optional.
Useful background:
EVOLVING_SURFACE.mdandGOVERNANCE.md.How to start
Reply in the public Urusilla Discussions
with these six short fields:
The maintainer will answer publicly with
accept,scope-correction, ordeclineand a reason before substantial work begins. The first sprint stays ina public issue or pull request. Do not provide credentials, private prompts,
private conversations, employer-confidential data, or legal identity documents.
After one useful public sprint, both sides may decide whether to continue as an
ongoing research pair or small team. A synchronous meeting is optional and
requires a separate mutual decision; it is not a condition for technical
credit.
Credit, rights, and authority
Accepted favorable, unfavorable, and null evidence receives equal attribution.
Contributors retain copyright in their contributions unless a separate written
agreement says otherwise; included contributions are licensed under Apache-2.0.
Collaboration does not automatically grant maintainer, release, registry,
signing, account, domain, or treasury authority. Any official role requires a
separate, explicit, bounded public delegation under
GOVERNANCE.md.No payment, token, employment, equity, governance vote, or future reward is
promised. The first sprint should use no paid calls unless the contributor
independently chooses and clearly discloses them.
What success looks like
The first success is not a star count or an endorsement. It is one independently
authored artifact that changes a test, narrows a claim, reveals a boundary, or
produces a reproducible disagreement. A continued collaboration should then own
one public question end to end: preregistration, implementation, evaluation,
unfavorable-result preservation, and a short written conclusion.
All reactions