-
Notifications
You must be signed in to change notification settings - Fork 0
Where This Is Going
Three tiers, and the difference between them is commitment rather than ambition. The first is funded or written. The second is designed and waiting on one decision each. The third is speculation offered for argument.
Committed work: this year into 2027.
A grant proposal is under review, with a decision expected in October: 2027 workshops where student teams compile a scenario of their own region's grid from public system-operator records. Their scenarios enter the library with their names on them, because named and citable authorship is the whole incentive.
The teaching doctrine is one line. Delivering our scenario about our grid is charity. Teaching a team to compile theirs is capability.
Three papers are written and under review. The instrument itself. A compiled Winter Storm Uri scenario. A compiled Iberian Peninsula scenario.
The claim they stake is that scenario compilation from public data, certified optima, and replayable scoring are a repeatable method rather than a one-off build. Anyone can then hold the platform to it.
Winter Storm Uri, compiled: a multi-day cold snap calibrated so that a current gold holder should not medal on the first attempt.
The Peninsula, the marquee. A reduced Iberian model anchored on the April 28, 2025 collapse, in two acts: survive it, then be useful during restoration. The scenario players already know from the news.
The grid answers back. In one scenario today, your deviation moves the price you pay, your relief can visibly downgrade an emergency, and the settlement reports what your choices did to everyone else's energy cost. More days in that era are coming.
Designed, not committed. Each is one decision away.
The founding commitment was that a human at a browser and an agent at an API are two implementations of one interface. Organizers run the physics; competitors run the agents. That boundary gives reproducible scoring and resistance to cheating by construction.
Humans first was deliberate. The seat is built. The second kind of player has not arrived yet.
There is a live and largely evidence-free public argument about whether a large data center raises your electricity bill. It is being fought in city councils and dockets, mostly by assertion.
The answer depends on behavior, and behavior is the one thing this platform measures. A flat load riding through peak raises everyone's costs. A load that curtails, stores, and improves its load factor can lower them. Same facility, opposite outcomes. The engine already computes the first-order version of that number.
Today the seat is one facility, one day. The campaign widens it to a business lifetime: take a capital budget, build the plant, play the compiled days against what you built, settle, fund the next build cycle, play harder days.
Every flexibility rule is finally a rule about where capital goes. Storage, on-site generation, and permits against raw capacity. The campaign measures whether an operator makes that investment voluntarily, and at what cost of funds it stops being rational.
That is a number the standards conversation has never had.
A later phase turns the rules themselves into the game. The Docket: intervene in a proceeding, and the score is the rule you get. The Ballot: draft a ride-through standard while other players vote their interests. The Protocol: write the market rule, build the coalition, then play under the rule you wrote.
The standards corpus today is one American peak day at a time. A day from a system where load shedding is routine is not a nice-to-have. It is the evidence the corpus most conspicuously lacks, and the way to get it is to compile it with the local community, from their operator's public data, with their names on the result.
Four inventions. Speculation, offered for argument rather than as plans.
Today a draft requirement meets reality only after it ships. A proposed rule, a ride-through obligation or a curtailment-capability floor, can be compiled directly into a scenario within weeks of the draft.
Then a working group watches a hundred operators play under the proposed rule before the comment period closes. Where they comply. Where they exploit. Where two reasonable readings of the same sentence diverge.
Ballot comments then arrive with episode logs attached. Not "we believe this is ambiguous," but forty replays of the ambiguity being exploited.
Regression testing for standards. It needs one working group willing to try it once.
Every AI benchmark wants two properties and almost none have both: an absolute reference rather than a relative ranking, and reproducibility by construction. This platform ships both today, as certified optima and the replay contract.
So publish the field's number. Gap to certified optimum, per scenario, per agent. A research group's grid-interactive agent gets a score that means the same thing next year.
The difficulty ladder is already a benchmark suite. The contributed library is its growth path. The infrastructure exists. The naming ceremony does not, yet.
After every major grid event, the record becomes a report that few read and fewer can interrogate. The compiler makes a different artifact possible: the event, playable, while the docket is still open.
Within weeks: the day compiled from public data, provenance stating what is and is not known, a certified optimum where one can be proven, and every armchair critic invited to do better than the operators did.
The event's black box, replayable by anyone. Regulators, editorial boards, classrooms, juries of public opinion.
Nothing new needs inventing here. The discipline is speed.
A pilot may not fly a 787 on the strength of a résumé. They hold a type rating, earned and maintained in a simulator, against the scenarios chosen to be the ones that matter.
The energy desk of a gigawatt-class campus has no equivalent. The person deciding whether to ride through, shed, or island during the region's worst hour holds whatever their employer decided sufficed.
A scenario battery, certified optima, and replayable records are the apparatus a credential needs: demonstrated play, on the record, against the days that matter.
Continuing education with a leaderboard. Whether that road is worth scouting is an open question, and a real one.
The first tier is why the library needs scenarios. The rest is why it needs good ones.
Exedra
Why
How it is built
What you write
Boundaries
For reviewers