Repository navigation
association 4.4.0
-
The subject reading's list of ordinary words ships with the package (
query/words.txt.gz, from Debian'swamerican2020.12.07, with its SCOWL notice inwords.COPYRIGHT) instead of being read from/usr/share/dict/words. On a machine without that file, "Best record from 2010-11 to 2018-19" was read as a question about Travis Best: the fallback kept only the month abbreviations, so any word that is also a whole player name ("best", "head", "pointer") became a player. GitHub's runner is such a machine, which is why CI failed from7b8b7a2on while the same tests passed locally. The shipped list holds exactly the 64,005 words the system file gave, so answers on a machine that has the file do not change. A wheel missing it now raises instead of silently answering differently. -
The subject reading writes the first field:
subject.apply_subjectputs the players the question was read to be about intoplayer/players- a router name the question never held replaced by the question's own spare name, or refused by name - beforeoverride_invented_playersruns over the same field (which the golden then shows finding nothing to do). Each replacement is a recorded decision. A router name the question spells differently - a bare "Jokic" left bare, "Jaylen Tatum" for the question's "tatum", "Deron Williams" two edits from "derozan" - takes the question's own resolved name (Subject.named), which was_question_derived_player's job. -
Decisions as data:
query/decisions.py'sDecision(stage, field, before, after, reason)records what was decided on the way to an answer - a reading of the question, an override of a routed field, a default - kept on the history record (adecisions:section of JSON lines beside the trace), onAnswer.decisions, on the API'sAnswerResponse.decisions, and shown by the page under a "decisions" disclosure; never parsed back out of the trace. The subject reading is the first producer: every fast-path answer now records who the question was read to be about, beside the repair chain that still writes the slots. -
query/subject.py: one reading of who a question is about (Subject,read_subject,question_supports), made from the question's own spans with the router's slots as a hint only where the question supports them. Measured before it was written: agrees with the fast path's ten-step repair chain on 279 of 290 recorded questions, and is right on the four where the chain is wrong. Nothing reads it yet - the roadmap's plan item 1 wires it in step by step against the golden. -
team_outlook's card on the page no longer repeats its own boxes as notes (only the BPI's offense/defense split, the strength-of-schedule rank and the other snapshots remain beneath it), and a postseason snapshot's second record is labeled "regular season" (regular_season_wins/losses) rather than "projected" - ESPN's projection columns hold the finished regular season once the playoffs are on. -
"Vs southeast division", "against eastern conference teams", "in the west": the router keeps the conference or division's name with the word as the
situation, so the relation reads it (it captured the bare word "division", which the alignment reader refused by name - #213, yardstick-v2 F055). A bare word still captures and is still refused honestly. The question then reached the compiler as aplayer_splitsline over a boolean measure, which builtAVG(<boolean>)and crashed in DuckDB: a boolean measure on a one-figure intent is now its count (8 road double-doubles against the division, the key), andcore._aggrefuses to average or sum a condition at all rather than send it to the database. -
Saving an edited note again replaces the earlier one in the history file instead of appending a second:
POST /api/notesreturns the line it wrote, and a save passing it back asreplacesrewrites that line in place (append_note(replacing=...); a line no longer in the file appends, so a correction is never lost). The page's note box holds the line its last save wrote and passes it back, so editing a note and saving again is a correction, not a second note. -
"Who attempted the most three pointers" ranks attempts, not makes (the question's own word decides the column), and a championship question is refused before a team ranking answers it fluently and wrongly - titles are not on record as such.
-
The web page shows every rendered answer's sentence first and at reading
size (it was the muted 12.5px caption), shows what the sentence goes on to
say beneath the table - a footnote's key, a coverage caveat, how a name was
read - instead of folding it under the text toggle (data["notes"], filled
by the templates andagent.pyas they attach them; the compiler's notes
are lines of their own), draws the sparkline at full width, folds a
season-long game table under its sparkline, renders a table-less answer as
prose rather than monospace, formats a rate column as a percent, and draws a
composed answer by its skeleton (rows, grouped, scalar) whatever intent the
router filed.team_record's month tables,period_leaderboardand a
player log's venue, result and per-game line render too. -
scripts/preview_answers.pyrenders any recorded question through the page
without the router - from a live-run jsonl, a history directory or an
explicit case - and screenshots each answer in headless Chromium. -
The web page's tables are typeset in the page's own proportional font
(font-variant-numeric: tabular-numskeeps digits aligned) instead of
monospace, with a header rule and row separators; a percent column now
right-aligns like any other number. A chart-only answer (fingerprint,
shot_chart) shows a headline and its notes for the first time - it used to
render nothing above the iframe but the artifact itself, with which players,
which season and a substitution note ("only one of them matches anybody in
the warehouse") visible only under the collapsed text toggle. A player's
game log renders the extra columns a question asked for (data.columns,
e.g. FGM/FGA) instead of always the same four; a single-leader question
("who leads the league in total netpoints") still renders as a ranked table
instead of falling back to monospace text; a rate column whose label spells
out "percentage" or "pct" (a leaderboard's own wording) formats as one; a
margin/differential column keeps its sign;player_stat's per-game figure
and season total no longer share one label ("PTS" twice);period_split
folds a long per-game table under its sparkline the wayteam_quarter_points
already did, and a sparkline needs three points, not two, to draw. A
redundant caption (equal to, or a prefix of, the headline) no longer shows
twice, and a note re-derived from an answer's own sentence no longer repeats
underneath a body already built fromdata-data.notesalone once a
renderer has drawn something, sinceanswerNotes's text-parsing fallback
could only guess at the same sentence a structured body already read
correctly. A numeric table column now shows one decimal throughout once any
cell in it has one (table()'spad), except a synthesized summary row (a
game log's "per game" line,skipPaddingRows) or a transposed table where a
column mixes a count with an average (player_compare,player_matchup,
pad: false), which format their own decimals per row instead. -
A ranked table (leaderboard, threshold_count, period_leaderboard) shows its
top 10 rows and folds the rest under their own count instead of listing
every one - a "top 50" question drew all 50 rows regardless of what was
asked. -
player_historyfolds a season table over 12 seasons ("past 20 years") the
same way a long game log does, and a sparkline is only drawn from points
that are real finite numbers - a mismatched key between the data and what
the chart reads it under used to draw a line throughNaNcoordinates
(see ISSUES.md, "Template data the page cannot render from",
player_history/twoPointFieldGoalPct) rather than simply not drawing one. -
The web interface fits a phone. Below 600px the coverage pills are one row that scrolls sideways, and the header, footer and answer padding tighten - the chrome took about 80% of a 667px-tall screen and now takes under a quarter. A column heading wraps rather than pushing the next column off-screen (a comparison's second player started past the edge), and a table wider than its answer shows a shadow on the side with more to scroll. On a touch screen the question box is 16px, so iOS no longer zooms into it, and the keyboard hint is hidden; the page is sized to
100dvhand clears the notch and home indicator. Shot charts and fingerprints scale down to the frame they are drawn in (svg { max-width: 100% }) instead of being clipped, with the fingerprint's skill labels enlarged so they stay legible; at desktop width the page is pixel-identical. -
The web interface shows whether it is connected - a dot in the top bar, green when the server answers, amber after one missed check, red after three - and reloads itself when the server restarts (a new commit or a rebuilt warehouse, where the server is deployed to restart on those), but only when that loses nothing: with an answer on screen, a draft typed or a question in flight it says "updated · reload" instead, and a click reloads. The page polls the new
GET /api/ping(association.web.app.PingResponse): an instance id drawn fresh each time the server starts, and whether it is answering. Unlike/api/healthit opens no warehouse connection and never probes ollama. -
Answer.history_filenames the.history/record an answer was written to (set byAgent.askin the same step that writes it); the web runner reads it off theAnswerrather than parsing the[history] ...trace line, which the web path carries verbatim. -
The web interface can annotate an answer: a note control under each answer posts to the new
POST /api/notes, which appends a timestamped[note ...]line to that answer's own.history/file, guarded by name the same wayGET /api/artifacts/{name}is.association.query.history.append_notewrites the line (escaping an embedded newline so the file stays one thing per line), and the answer event now names its own history file (AnswerResponse.history_file,association.web.runner.Answered) so the page has something to attach a note to. More than one note per answer is allowed; each save appends rather than replacing the last. -
The compiler reads an ordinal season over everyone ("Most points in 15th season played": each player's own 15th regular season, on the league read), a team's name in the
playerslot as the team narrowing of a league-wide read ("oklahoma city thunder all-time triple doubles": 180, by player, the whole count behind a windowed list), and the router's fillerplayerword as no player. -
Three more shapes are refused fast with their cause instead of reaching the agent: a team's stat other than points by quarter (the linescore holds points only), bench points (derivable from the starter flag, read by nothing yet), and an attempts or minutes floor on a league ranking (the compiler now refuses it by name rather than declining).
-
The pair relation is on the relation:
player_matchupnarrows the first player's games through the shared step, so a teammate's absence ("curry vs lebron without kd"), a venue, a date, a starter half, box-score lines and a calendar are honored and stated in the heading;orderand an ordinal season are excluded with their reasons. A matchup with no meetings under a narrowing says the meetings are missing from those games, not that the two never met. -
A player filed as the
opponentis the second of two players, andplayer_matchupanswers their meetings ("lebron vs kawhi head to head", "jay huff game log vs embiid") - the refusal that said games between two named players were "not read yet" was false and is gone. -
A team's conference and division are read from the standings, and both
relations narrow games by an opponent's (K3-2, ISSUES.md #25). ESPN's
standings response, re-requested at&level=3, nests a division level the
plain request (standings's own source) does not carry - probed live
across the table's floor, the same two conferences hold four divisions
before the 2004-05 realignment and six after.fetch.parse.parse_team_alignment
reads it into a newteam_alignmenttable (season, team_id, conference,
division; same 1988 floor asstandings, since it is the same endpoint).
query.calendar.parse_alignmentreads asituationvalue naming one -
"vs the west", "against eastern conference teams", "vs southeast
division", "in the west" - beside the existing calendar reader
(parse_situation), andNarrowed.narrow_alignment/TeamNarrowed.narrow_alignment
narrow to games against an opponent aligned that way IN THAT GAME'S OWN
SEASON, applied by the same shared steps every othersituationshape
goes through (_apply_situation,templates/common.py) - so every
template that already honoredsituation(game_log,player_stat,
threshold_count,player_splits,record_when,streak,team_record,
team_leaderboard,team_quarter_points,team_stat, and more) reads it
at once.query.refusals._non_calendar_situationstops refusing it by
value and, where the words are right but the phrasing is not ("the Central
Division these days"), says so rather than the generic "not read" message. -
"luka td3s home" counts his triple-doubles at home (48, through the compiler) instead of going to the agent: "td3s" is no longer read as a period word, and
route()files it asstat: triple_double- never a shot value of 3 - with a named player'sotherbecomingplayer_stat(#194). -
"Kd vs clippers 2h at home gamelog" is Durant's second halves against the Clippers at home (
period_split), not the Clippers' own half: a player and a team filed together inplayerson a "vs" question are read as the player and his opponent. -
"Best NBA record" (or "best league record") with no team named routes to
team_leaderboard, as "best record" already did - "Best NBA record since January 31st 201" came back asteam_recordwith no team and fell through. -
A team-only question naming one player and a team the question never mentions is refused by the player's name: "towns home rec including playoffs since 1/26/20 vs spurs" routed
team_recordwith an inventedteam='Toronto Raptors'and would have answered the Raptors' record. A real team counts only when the question holds a trace of it. -
"bam adebayo career games in the month of march" is his game log in March (118 games) rather than a fall-through: a
threshold_countover the line 0 on the stat "games" isgame_log, with the subject restored, and "in the month of " is read as the month like "in ". -
"Since 2000-01" (and "since 2000-2001") starts at the season ending in the later year, 2001;
route()read only the leading "2000" and started a season early (#207). -
A player name the router invented in
opponentis checked against the question like the subject is (#206): "jay huff game log vs Embiid" routedopponent='Nikola Jokic'and the refusal named Jokic; the one player the question names that the subject does not claim now takes its place, and with no exact replacement the name is dropped. A team opponent is untouched. -
A composed league-wide log names the player on each row ("every player ... top 10 by points" listed dates and figures and never said whose they were).
-
The compiler treats
ranked_byandteam_restoredas its own slots (COMPILER_SLOTS) rather than refusing them as scoping the relation cannot honor - measured live, "players with the highest scoring triple doubles" reached the compiler and fell through on the marker alone. -
"Players with the highest scoring triple doubles" ranks the triple-double games by points instead of counting them:
route()filesranked_by, a slot no template honors, soleaderboardrefuses and the compiler's boolean-game ranking answers (the bare count "most triple doubles" is unchanged). -
A team's half named only by its nickname ("Celtics 2nd half scoring this season") routes to
team_quarter_pointswhere the model filed noteamslot and the question went to the agent. -
What nothing here can answer is refused fast, with its cause, instead of
handed to the slow agent:association.query.refusalsis the step after
the template and the compiler both decline - a playoff round (the games
carry no round label), an age (no birth dates on record), a conference or
division named as a team (team_record'sopponent, still nothing to
tally), a stat other than points by quarter, a game log "vs" another
player, a team where a player belongs. Measured on the yardstick's
fall-throughs, each took 30-120 seconds to reach an agent answer that was
wrong or never came. (A conference or division named as an opponent's
situation - "vs the west" - now answers instead; K3-2 below.) -
"Stats for the sixers when maxey scored 20+ points" is the team's record
under the condition (record_when), not the player's own average - two
readers agree before the intent moves; and "for " is a player's
tenure only when the player is named before it. -
A
period_splitwindow the question never named is dropped ("each game",
"games": a fillerlimit: 1printed one row under a whole-season total),
and anopponentthat only repeats thewithoutlist is not an opponent. -
query.composenow carries the box-score caveats a template's own
answer would (#197, ISSUES.md, closing the entry's remaining half - the
coverage-floor half was fixed earlier the same week). A composed
answer over box scores never said which teammate's absence it read as
"without", how many games it left out for an empty ESPN box score, how
many of its games were rebuilt from play-by-play rather than fetched, or
that a career predates box scores entirely - all of it
templates.common._box_score_notes, which the six relation templates
already call andcomposenever did.core.run()now calls it itself
(a named player only - the league-wide subject has no ONE player's career
to check a floor against) and returns the result asnotes, which
answer()appends to the sentence the same way it already appends the
coverage caveat. The scratchrebuilt_showncolumn
core._scalar_selectsadds to compute the rebuilt-line count for a
scalar/groupedread is popped back off before the rows reach a
caller, for every subject - a named player's own reads and the
league-wide one alike, since it is an internal detail of how the count
was gotten, not a value either was asked for. No relation counterpart
exists for the team subject (compose/team.pyreads
games/team_season_stats, never a player's box score), so nothing was
added there. -
A position word in the router's own
playerslot ("shooting guard") is
read as the position-group subject rather than a player name nothing
resolves to (F056, ISSUES.md #160).query/compose/move.py's new
_position_only_player/_drop_position_only_playerclearsplayerwhen
it holds NOTHING but a position word, before the rest of the module's
existing position handling (_position, reading the question text) takes
over - so "highest 3 point percentage ... by a shooting guard" no longer
tries to resolve "shooting guard" as a name. A "with at least N games"
phrase now replaces a league ranking's default minimum sample
(_ranking_minimum); a unit the relation cannot apply as aHAVING
clause (attempts, minutes) is refused by name rather than silently
dropped or misread as a games count. Fixed the same session: "at least"
contains the word "least" as a whole word, which flipped a "highest ..."
ranking to ascending order whenever a minimum-sample phrase was present -
measured against the real warehouse before the fix ("highest 3-point
percentage ... with at least 40 games" answered lowest-first) and pinned
by a fixture test after it. Warehouse-verified (season 2025, which
carries specific position codes - see DATA.md): "highest 3-point
percentage in 2025 by a shooting guard with at least 40 games" correctly
ranks Alec Burks (42.5%, 49 games) first, descending. Found and recorded,
not built here: an attempts floor is still refused (noHAVINGclause
for one yet), and the CURRENT season's roster data mostly lacks specific
position codes at all (DATA.md), so a position-group ranking for "this
season" alone still returns nothing even once the subject reads
correctly. -
A composed league-wide
threshold_countnaming SEVERAL " "
lines at once ("33 points and 13 rebounds and 10 assists 2 blocks and 2
steals", F161) now lists the games clearing every line, rather than
refusing for want of a single line to count.query/compose/move.py's
_numbered_stat_linesreads every " " pair straight out of the
question text (the sameMEASURE_WORDSlookup a single threshold already
used), and_everyone_multi_line_gamesanswers rows over everyone under
all of them as predicates, naming who had each game - not
_everyone_threshold_count's per-player COUNT, which still answers a
single line exactly as before. Warehouse-verified: over 1994-onward
regular-season box scores, exactly 12 games clear 33+ points, 13+
rebounds, 10+ assists, 2+ blocks and 2+ steals at once (direct SQL
cross-check against the same five columns agrees). Found and filed, not
fixed here: a single-line count still falls through when the router's own
statalready names the phrase's own column (_everyone_threshold_predicates's
ranking-measure dedup applies where it should not), reproduced live
againstnba.duckdb(ISSUES.md). -
A composed league-wide ranking of a boolean measure ("highest scoring
triple doubles") ranks the qualifying GAMES by another measure, rather than
counting them (#199, ISSUES.md).query/compose/move.py's
_everyone_boolean_game_rankingreads "highest"/"biggest"/"most in
a"/"top" over a boolean predicate (triple_double,double_double,
fouled_out) as rows over everyone, ordered by the question's own stat
word (points by default), not_everyone_ranking's per-player average -
which would have neededminimum_gamesqualifying games just to rank
anyone. "Ever"/"all-time" now moves the otherwise-current-season default to
a career read (_everyone_career_slots), stated because a league-wide read
with no season named defaults to the current season. Warehouse-verified:
"biggest triple double ever" returns Nikola Jokic's 61-10-10 (2025-04-01 vs
Minnesota), the regular-season high the router's own corpus keys against.
Reachable only onceleaderboarditself refuses this framing instead of
answering a triple-double COUNT first - the router/template half of #199,
not built in this worktree. -
team_record's combined-season-types sentence names each half's own
starting season (#204, ISSUES.md). "Warriors all-time record including
playoff record at away" used to carry only "Note:" lines from its two
halves, leaving the reader unable to see that the regular-season half
starts five seasons after the playoff half's own floor. The sentence now
reads "(523-791 regular season from 1993-94, 51-52 playoffs from 1989)" -
each half's first season, read from its owndata(_games_recordnow
carriesfirst_season/last_season, computed fromteam_games's own
season column for a single named season and from a small MIN/MAX query
otherwise) rather than parsed out of either half's sentence.
A postseason's first season is read by the calendar year it was actually
played (year(tg.eastern_date)), neverteam_games.seasondirectly -
that column is ESPN's own pre-1993-94 label, which names a season by the
year it STARTED, not the year its games were played (the same fault
templates/splits.py's_team_season_rangealready guards against).
Measured directly against the warehouse before trusting the constant
"1989" this file's own prose used elsewhere: the Warriors' own earliest
playoff game reads 1988 by the raw label and 1989 by calendar year -
confirming the raw label would have been wrong here too.
Warehouse-verified against the Warriors' road record. -
leaderboardcan show each player's team beside their name (F017,
ISSUES.md). "Who are the top 50 in total adjusted netpoints with the
team they play for" asked for a team beside every row and got none. A
fieldslist now accepts"team"alongside the existing box-score
columns (_leaderboard_fields); it adds a "team" column read from
player_game_log- the team a player played his most recent game for
that season and season type, so a mid-season trade shows the last team,
said once in a note ("Team is each player's most recent team that
season.") rather than per row. Player only, and only where the ranked
metric has a season type to look the team up by.run_leaderboard's rows
now also carry each row's athlete id on the side
(LeaderboardResult.athlete_ids, aligned by index, never a key inside a
row dict) so a caller can resolve something the ranked table itself does
not carry, without that id riding along intotoolbox.get_leaderboard's
JSON for the model to read. Warehouse-verified: a 50-row NetPoints-per-100
ranking withfields: ["team"]now lists each player's team, all 50 rows
present (the row LIMIT was already honored correctly before this change -
confirmed directly against the warehouse; F017's own capture was cut off
by the yardstick log's own display truncation, not a real limit). The
router does not yet emitfields: ["team"]for wording like "with the
team they play for" -ROUTER_SCHEMA'sfieldsenum has no"team"
entry - so this is the template-side half of the fix; see ISSUES.md for
the remaining router-side gap. -
player_splitsadds a column for a stat the standard line does not
carry, or refuses naming it (F159, ISSUES.md). "Quentin Grimes
individual gamelog usage rating without joel embiid" showed the standard
split columns (G, W-L, MIN, PTS, REB, AST, STL, BLK, TOV, 3PM, FG%),
which have no usage-rate column at all, so the actually-asked-for stat
was simply absent from an otherwise-correct read. A namedstatthe
table already carries changes nothing (unaffected); one in the new
SPLIT_EXTRA_STATSmap (today,usage_pct) reads straight off the
relation and gets its own column - "USG%" - read directly from
player_game_log, per-split, the same as every other column in the
line; any other named stat refuses by name rather than answering without
it. Player only - a team split has no per-player rate to show, and a team
subject now refuses the same stats a player subject can show a column
for. Warehouse-verified: Quentin Grimes's splits without Joel Embiid now
carry a USG% column with real per-split values. -
player_statprints a made-count stat's attempts and percentage beside
its total (F051, ISSUES.md). "Davion Mitchell 3 point stats" (a single
named made-count stat: 3PM, FGM or FTM) used to print makes and games and
nothing else. It now reads the attempted sibling column alongside the
made one - the same "out of how many?" disciplineSHOOTING_STATS
already keeps for a bare percentage - and says "That is 90 of 228
(39.5%)." instead of "That is 90 in total.", over a season, a career and
a box-score-narrowed read alike. Never for a multi-stat line (nostat
named), where singling out one entry's attempts would read as though only
it needed the qualifier. Warehouse-verified against Davion Mitchell's
season (90/228, 39.5%) and career (395/1,114, 35.5%) three-point lines. -
player_historystates the combined career figure for a career span, not
only the season-by-season table it sums to (F041, ISSUES.md). "Show me
sga's career 2pt percentage" printed a table whose rows summed exactly to
55.1% (3,945 of 7,166) without ever stating it. A career span now adds a
line under the table: the games-weighted total (makes and attempts summed
across every season shown, never a mean of means) for a shooting
percentage, and the plain career total (the stored season total, or
avg * gameswhere the table has none) for a counting stat. Warehouse-
verified against Shai Gilgeous-Alexander's career 2PT%. -
game_lognames how many games a window cut from, not only how many it
shows (F149, ISSUES.md). "Mikal Bridges game log with less than 15 FGA
and less than 35 minutes" cut 39 qualifying regular-season games to the
default 10 and headed the answer "last 10 games", with no word about the
29 left out. The heading now says "last 10 of 39 games" whenever the
window (a default limit or an explicit one) actually cuts the narrowed
games, read withaggregate_sqlovercommon.whole_span- the same
shared steps every other count over the relation uses, never a
hand-writtenCOUNT(*). Warehouse-verified against the traced question. -
A composed answer over a closed season range names the range it counted -
"regular season career (2020-2022)" - where it said "(2020 on)": the count
already stopped atuntil, so the sentence stated a scope the number did
not have (yardstick-v2 F036). A both-season-types span is named as such
(compose/sentence.py). -
team_recordhonorsseason_type_unstated, combining both season types
instead of silently answering one (F116, ISSUES.md). "Warriors all-time
record including playoff record at away" used to readseason_type: 3and
name only the postseason, dropping the regular season the question also
asked for._team_record_combined_typesreads each type through
team_record's own existing single-type routing and sums the two records,
naming each component - "574-843 combined on the road, including the
playoffs (523-791 regular season, 51-52 playoffs)" for the Warriors,
warehouse-verified. Measuring this against the yardstick key's own 601-865
(550-811 regular season, 51-54 playoffs) found the key's figures come from
an UNDEDUPED count ofreal_games/games: the 1994 postseason has the same
phantom double-labeling under bothseason1993 and 1994 that the regular
season is already documented for (DATA.md, "Season 1993 is a phantom"),
and it is not unique to the Warriors - 77 of the 1994 postseason's own
events carry both labels.team_games's existingQUALIFYdedup already
collapses this correctly; the key's undeduped total was not corrected to
match, since doing so would reintroduce the double-count this relation
exists to prevent. -
team_move_pointno longer treats the router's"any_team"placeholder
as a real name. Found re-running the K2 golden harness after landing the
team subject: "rebounds allowed per team" filesteam: "any_team", which
_resolved_teamRAISES for (unlike an unmatched name, which returns a
clarification) - so the question reachedrun_teamonly to fail there,
through a noisier path than the decline_everyone_point's own guard
already gives it ("allowed").team_move_pointnow falls back to scanning
the question's own words the same way it does for a droppedteamslot. -
compose.answerchecks the coverage floor and carries a partial-season
caveat (#197, ISSUES.md). It used to answer a season under a table's
floor as confidently as a modern one - a 1990 box-score question read "0
games" rather than refusing, which looks like a real zero and is actually
no data at all. Both subjects now call the same two functions every
relation template does (check_coverage/coverage_caveat), the team
subject through its owncompose/team.pycounterparts
(team_coverage_refusal/_team_coverage_note), warehouse-verified against
ESPN's real 2001-postseason gap note on both a player and a team question.
The box-score caveats (_box_score_notes- empty box scores, rebuilt
lines, a career predating box scores) remain open, filed in the same entry. -
A team's own total or differential is stated, not just its games
(F128/F129, ISSUES.md). "Total points scored by the Toronto Raptors in
the last 10 games" and "Knicks point differential over the last 7 games"
both narrowed to the right games already but never said the number the
question actually asked for - a wrong-shape answer wearing a right one's
clothes.game_log's team half now states it (_team_game_log_total_line,
read fromstat), in both the single-season-type and the mixed-season-type
("last N games" naming no type) readers. Warehouse-verified against both
targets exactly: 1,130 total points (3 regular-season games plus a 7-game
playoff series) and a +62 (+8.86/game) differential over the Knicks' Finals
run. -
The team as a subject in
compose(step 3, K1).compose.team-
TeamQuery/TeamResult/run_team,team_move_pointinmove.py,
team_sentenceinsentence.py- answers a question whose grammatical
subject is a team and names no player, over two readers: an unnarrowed
season TOTAL straight fromteam_season_stats(not the per-game average
team_statalready gives - "how many 3-pointers have the Magic made this
season" is 961, warehouse-verified exactly against ESPN's own total, not
team_stat's 11.7/game), and a narrowed sum over the team-games relation's
own points/points-allowed/differential columns for a single season type.
team_named_inrestores a team slot the router drops entirely, the same
repairplayers_named_inalready makes for a dropped player - and is tried
BEFORE that player repair, because "magic" is also Magic Johnson's given
name and the player repair would otherwise invent him from a team
reference. A box-score count (not a game-outcome figure) narrowed to a
window refuses rather than silently answering the season, since the
relation has no box-score join yet (filed in ISSUES.md). Not yet reachable
from the live pipeline for every shape it answers -compose.answeronly
runs after a template refuses, andleaderboard/team_statdo not refuse
a team-shaped question today - so F127's fix ships as tested compose
infrastructure while F128/F129 ship live, throughgame_log, above. -
A since-bounded team leaderboard now counts every franchise, not only
the ones with a game in the span (F100, ISSUES.md). "NBA team with
least playoff wins since 2022" ranked "of 28 teams", silently dropping the
Charlotte Hornets and Washington Wizards - neither made the playoffs in
that span - from a ranking about exactly that._team_leaderboard_since_records
now LEFT JOINs fromteamsinstead of inner-joining the games it finds, so
a team with none reads 0-0 and ranks among the "fewest wins" it genuinely
tied for, rather than being left off entirely. -
The team-games relation reaches parity with the player one:
situation
anduntil(step 3, K1).TeamNarrowed.narrow_calendarmirrors
player_games.Narrowed's own - a weekday, a month, a fixed holiday, or
"since " within each game's own season - applied by the shared
templates.common.team_gamesstep and reached "for free" by every template
that calls it (team_quarter_points;head_to_head's calendar reading is
filed as a follow-up, not built this round).team_recordgains the same
narrowing directly (it does not call the shared step for its own
season-record path), replacing its old bare-month-only reading - "the
Knicks' record on Christmas" and "... on Saturdays" now answer, warehouse-
verified (2025-12-25: Knicks 126-124 over Cleveland at home; three home
Saturdays this season, 1-2).untilis the inclusive last season of a
since-bounded range (a decade, or "2019-20 to 2023-24"), read the same way
sincealready is -team_record,team_leaderboardandhead_to_head
all honor it, andteam_leaderboard's span-bounded record now answers
"best record from 2010-11 to 2018-19" (San Antonio 509-213, ahead of Golden
State 479-243 and Oklahoma City 465-257 - matches the warehouse exactly).
team_record's by-month split also gains asince/until-bounded form,
one table per season rather than refusing outright - "Knicks record by
month 2024 2025" now answers both seasons' tables (warehouse-verified
against ESPN's own 50-32/51-31 season totals, month for month). -
A team the question names as its own subject, dropped by the router on a
leaderboard- orteam_stat-shaped question, is restored instead of
silently ranking or refusing about the wrong thing. "how many 3 pointers
have the magic made so far this season" routed toleaderboardwith
stat/seasononly - noteamat all - and ranked the league's
individual leaders in makes, the Magic never named (yardstick-v2 F127).
entities.teams_named_in(the team counterpart ofplayers_named_in)
finds the one team the question names, and
entities._scope_from_question_team_subjectrestores it intoteamso
query.composecan see it. Forleaderboardalone the restored value is
also markedteam_restored- a slot noHONORED_SCOPINGentry lists, so
check_scoperefuses and hands the question tocomposeinstead of
leaderboardquietly ranking players "on" a team that was meant to be the
whole subject;leaderboard's own, router-suppliedteamreading ("Top 5
scorers on the Lakers?") is untouched, since only a team THIS restore
itself wrote carries the marker.team_statgets no marker - an empty
teamthere already raisesTemplateUnsupported("no team named")on its
own, so restoring it is a strict improvement.teams_named_inexcludes any
span shorter than three letters and two measured common-word collisions
with a real team abbreviation ("was" -> Washington Wizards, "min" ->
Minnesota Timberwolves), the same disciplineplayers_named_inalready
applies to a player's name. -
A "last N games" period-split question naming no season crosses into an
earlier one when the current season has nothing. "zach collins first
quarter stats last 5 games as a starter" answered "No 2026 regular season
games found ... as a starter" - true of the box scores read and about the
wrong year, since his real last 5 starts are all from the season before.
games._period_split_cross_season_redirectretries the same narrowing
over his whole career, windowed to the newest N by date - the reading a
barelimitalready gets everywhere else on the player relation - and
answers with that season's own numbers and accuracy caveat when the found
games land in exactly one season; a question naming a season outright, or
a window landing across more than one, is unaffected and keeps the plain
refusal. -
A question naming exactly one player and no team, routed to a
team-only intent, refuses instead of answering the league's numbers.
"alperen şengün alltime record" routed toteam_leaderboard- no player
slot exists on that intent, and no team slot was filled either - and
answered the league standings, entirely off Sengun.entities.player_named_on_a_team_only_question
runs beforetemplates.common.TEAM_ONLY_INTENTS(team_record,
team_leaderboard,team_stat,team_outlook) reach their template,
naming the player in the refusal instead. Diacritics are already folded
(entities._fold); an inventedteamslot holding the player's OWN name
(measured live, a second run of the same question) is treated the same as
no team at all (entities._has_a_real_team), so the router's
nondeterminism on this exact question does not slip past the check either
way. -
A team named "for "/"with the " beside a player is restored
as his own tenure, and defaults the span to his career. "lebron stats as
a starter for Miami" answered his current (Lakers) season, "Miami" never
read at all - the player relation had no way to narrow to a team he no
longer plays for.entities.scope_from_questiontakes a new
restore_teamflag (templates.common.OWN_TEAM_RESTORABLE_INTENTS, only
player_stat), writing to a newown_teamslot - deliberately not the
router's ownteam, which a recorded golden case shows sitting as noise
beside an already-correctopponent(his own current team, redundant);
reading it directly would have silently renarrowed that case's 9 real
meetings down to 4.templates.common._narrow_player_gamesgained a
teamparameter (as opposed toopponent) for the actual clause
(pgl.team_id = ?). Corpus-measured against the same 380 questions as the
restore above: one true positive ("westbrook stats as a starter for
kings"), zero false positives. -
A player named with no scoring verb or possessive is restored for
single_game_high/threshold_count. "kawhi most threes in a game"
answered the league's single-game leaders, Kawhi Leonard's own 7 never
mentioned -router._SUBJECT_OF_HIGHneeds "kawhi scored" or "kawhi's",
and this shape has neither.entities.scope_from_questiontakes a new
restore_subjectflag
(templates.common.SUBJECT_RESTORABLE_INTENTS), corpus-measured
(scripts/check_routing.py's cases plus the StatMuse feed, 380 questions)
before shipping: 5 false-positive candidates turned up in the whole
corpus, none on these two intents, and the two recurring words behind them
("best" is Travis Best, "head" is Luther Head) are excluded going forward
byentities._named_only_by_a_common_word. -
A consecutive hyphenated year pair ("the 2023-2024 season") is one
season, not a range.router._RANGE_HYPHEN_YEARS, added earlier in this
cycle to read "2020-2024"-style ranges, matched ANY four-digit hyphenated
pair and so also matched a single season spelled out in full - silently
overwritingseason_text._SPAN's already-correct single-season read
(since=2023/until=2024 in place of season=2024). Now only a
NON-consecutive pair reads as a range ("how many 20+ point games did SGA
have 2024-2026?"); a consecutive one is left to the existing single-season
reader, exactly as "2023-24" already means. -
"since he joined the league" is a career span. "Show me luka's avg
assists since he joined the league" carried no_SPAN_WORDSmatch at all
("career", "all-time", "ever", "in/of history") and answered one season
(8.8 apg, 2019-20) where his whole career (8.23 apg, 514 games, 2019-2026)
was asked for.router._SPAN_JOINED_LEAGUE_WORDSreads it, anchored on
"the league" so "since he joined the Warriors" - a team question - is
unaffected. -
A near spelling in
withoutis taken rather than asked about, and the
shot-distance leaderboard refusal names the real cause. "de'aaron fox vs
magic last five games without wembyanama" refused "did you mean Victor
Wembanyama?" over a typo the question's own words resolve cleanly - the
true reason the question falls short of five games is a game count (one
qualifying game), not a name that failed to resolve._resolved_teammate
now applies the same visible-and-correctable default
entities.resolve_playeralready gives a bare surname
(entities.note_typo_reading) tosuggest_players' own single-candidate
result. Separately,leaderboard's shot-distance refusal said "no
leaderboard ranks shot distance", which reads as impossible and is false -
the key computes a real league leader straight fromshot_chart; it now
says the ranking is not built (filed inISSUES.md). -
"including the playoffs" no longer reads as "the playoffs only".
router._validate_season_typeconsulted_PLAYOFF_WORDSalone, and
"including playoffs"/"regular season and playoffs"/"playoffs included"
contain the word "playoffs", so a question asking to keep BOTH season types
silently dropped the regular season - "Payton Pritchard stats vs 76ers at
home including playoffs" answered 7 playoff meetings and then falsely told
the reader "Only 7 games ... in his box scores", never mentioning the 9
regular-season ones it never read; "jamal murray games with 2 threes
including playoffs" answered 3 (his 2026 postseason alone) instead of 59
(2026 regular season and postseason combined).router._BOTH_SEASON_TYPES_WORDS
now reads the phrase and setsseason_type_unstated(the flaggame_log's
"last N games" reader already carried), honored bygame_log,
player_statandthreshold_countthrough one combined
player_games.season_type_clause/BOTH_SEASON_TYPESread on the player
relation rather than a merge - simpler thangame_log's own row-interleave,
since an aggregate has no rows to interleave. A template that does not
honor it (team_record, on the yardstick's team-side example) is refused
bycheck_scoperather than silently answering the postseason alone. -
A closed season range ("2019-20 to 2023-24") no longer reads through the
present.router._validate_rangeread only an open "since 2020" or a
decade; every other range form ("2019-20 to 2023-24", "from 2010-11 to
2018-19", "between 2020 and 2024", "2020-2024", two adjacent bare years)
fell back to the model's own single-season slot, so "Portis vs bulls
2019-20 to 2023-24" answered 3 games of one season where 16 regular-season
games across five were asked for. Worse, even wheresinceWAS read
correctly,until- the range's other end - was declared nowhere and
honored nowhere, so a CLOSED range read as an OPEN one: this project's own
worst-failure-shape example, an unfilleduntilanswering "the 2010s" as
2010 through now._Span/player_games.Narrowednow carryuntilbeside
since, applied in one clause (_Span.clause) the same waysincealready
was, and every player-relation template that honorssincehonorsuntil
beside it (test_until_is_declared_wherever_since_is) -game_log's team
branch, which had silently droppedsinceitself for the same reason,
picked up both in the same fix. -
A composed answer names a position group in its heading, and prints
TS%/eFG%/usage as percentages. The first live run of the landed compiler
headed a log of centers "every player" (the filter was applied; the heading
did not say so - the value used was hidden) and printed the view's
fractional TS% as "0.5%". Both are the sentence, not the numbers. -
One compiler over the player-games relation (the skeleton spike, landed).
association.query.compose-core.Query/compile_query/run,adapt.to_query,
move.move_point,sentence.sentence, andanswer()in__init__.py- is the
step between a template's refusal and the slower SQL-writing agent, over the
same relation the six relation templates (game_log,player_stat,
threshold_count,single_game_high,player_splits,record_when) already
read. AQuerynames a point (skeleton, measures, aggregate, group, window);
to_querygives an intent's default point;move_pointmoves it with the
question's own words - a measure beyond a template's list, "most ... in a
game" as rows by measure, "how many ... won" as a career count, a league-wide
read with no player named (grouped by player for a ranking, rows for a
position group), and the number in a threshold phrase naming its own column
over the router'sstat. The relation supplies the narrowing through a new
shared step,templates.common.league_games- the league-wide counterpart of
scoped_games, narrowed by opponent, venue, team, box-score lines, calendar
situationand a position - so the compiler never writes its own clause on
the relation (test_templates_on_the_relation_do_not_narrow_it_themselvesnow
walks the compose package's source for the same forbidden tokens the six
templates are checked against). Measured against the real warehouse: 237
agree / 0 disagree over 387 recorded questions the six templates already
answer (232/0 at the spike's own master commit;situationlanding on the
relation since then answers more of them on both sides), 8 of 31 template
fall-throughs answered correctly, and the yardstick's K3 tally unchanged
(138 answered live, 3 by a template, 8 by the compiler, 26 still refused).
Fixed two bugs the spike's own code carried, both exposed only because
situationlanded on the relation after the spike measured 232/0: ateam
slot beside the player used to narrow a per-game average (player_stat's
shape) to his games for that team, when the real template reads no such
slot at all and silently returned zero games for a team he never played
for; and a bare non-ISOdateslot (the router sometimes files a weekday
word there, meant forsituation) was handed to the relation unvalidated,
raising instead of being read as absent. Not wired into the pipeline yet -
agent.py/answer.pyare a separate change - so nothing here changes what
a live question answers today. -
A template's refusal composes before it falls through.
agent.pyused
to treatTemplateUnsupported- raised bycheck_scopeor by a template
refusing its own slots - as the end of the fast path: log why, and hand the
question to the slow SQL-writing agent. It now offers the new
association.query.compose.answer(ctx, intent, slots, question)one more
try first, on the same point on the relation the template could not narrow
to. ATemplateResultback is answered exactly like a template's own -
answered_by="fast", the intent kept, the same name-reading and
coverage-caveat attachment, a trace line naming what was composed - and that
includes a refusal the compiler hands back (a clarification, a "no match"):
it looked at the question and had something to say, which is an answer, not
a fall-through.Nonestill falls through exactly as before this step
existed. This lands the call site and a contract stub
(association.query.compose,answer()always returningNone) only; the
compiler itself - the six-template golden the skeleton spike measured at 232
agree / 0 disagree - is a separate, parallel piece of work that replaces the
stub module wholesale. -
The player-games relation honors
situationwhere it names the calendar
(step 3, K3). "Garland on Mondays", "LeBron's line vs the Jazz on
Tuesdays", "his games in October", "on Christmas", "since January 31st" -
the router files all of these in onesituationslot, and every template
on the relation refused the slot outright (check_scope).query/calendar.py
reads the four shapes a game table can filter on - a weekday, a month, a
fixed-date holiday, and every game from a day of the season on (in the
calendar year that day falls in for each game's own season) - and
Narrowed.narrow_calendarapplies them over the game's US Eastern day, so
"on Tuesdays" is the night it was played. Declared once in
RELATION_SCOPING, sogame_log,player_stat,period_split,
player_splits,record_whenandstreakall read it, and each answer
says it ("on Tuesdays" in the heading). Asituationthat names anything
else - an age ("as an 18 year old"), "since returning" - is still refused,
now by value and with the shapes that are read named in the message; it is
never dropped. A conference or division is read too, as of K3-2 below. Of
the 19 distinctsituationvalues across every recorded corpus, 11 are one
of the four calendar shapes. -
Two router post-processing fixes, both slots the question does not
support arriving in the wrong place. A quarter or half question that
names a player but whose model reply dropsplayerentirely (filling
team/opponentinstead) used to read as the "team's own half" shape and
route toteam_quarter_points, a template with no player column - "How
many points did Jokic score in the 3rd quarter against Boston?" answered
(or tried to answer) Boston's own quarter.route()now recovers the
player from the question's own grammar the same waythreshold_countand
single_game_highalready do for their own dropped subject
(_subject_named_in), filtered against_is_team_nameso a genuine team
subject ("did the 76ers score") is never misread as a name; the surviving
team-shaped slot lands inopponentonly where the question's own words
say it (_team_slot_named_in_text), closing the ISSUES.md entry that
tracked it as a standing regression. Separately,team_leaderboardrefused
"nba team with least playoff wins since 2022" with "no team matching
'least'" - the router had correctly read "least" intorankand then,
independently, filed the same word intoteam, where no franchise is
named "least".route()now drops ateamvalue that is itself one of
theRANK_WORDStherankslot is read from, the same shape
entities.override_invented_playersalready applies to an invented player
name. Neither fix touchesROUTER_PROMPTorROUTER_SCHEMA(hashed
unchanged before and after). -
team_recordhonorssinceandgame_n;head_to_headhonorssince
andspan"career" (step 3, team cells). Both used to exclude these
cells fromTEAM_RELATION_SCOPING_EXCLUDEDwith a reason about the code
("not built yet"), which the declaration's own rule forbids - closes the
ISSUES.md entry filed when C4b'sTeamNarrowedgrewsinceand
series_gamebut neither template read them. "Celtics record since 2022",
"Lakers record in game 1 of each series", "Celtics vs Knicks since 2020"
and "all-time Celtics vs Knicks" used to refuse and fall through to the
agent; they answer now.team_record'ssincereads the same since-bounded_Spana whole career
already does (_record_narrowed, extended to build it through
common._span_of/common._team_span_clauseinstead of
team_metrics.games_scope, which only ever means "one season" or "every
season");game_nis the relation's ownTeamNarrowed.narrow_series_game,
the same checkcommon.team_gamesmakes for every other team template (a
series has games 1-7, a regular season has nothing "game 4" names). Both
compose with an opponent, a venue and a calendar month exactly as a single
season already does, and the header says so -game_n's "in game N of
the/each series" throughTeamNarrowed.filters(opponent=False), the same
idiomteam_quarter_pointsuses, rather than a hand-written phrase of its
own. Neither combines withsplit"month" yet, which refuses rather than
silently answering the plain month split instead.head_to_head'ssince/spanread throughcommon.team_gamesthe same
way its existing single-season path already does - new
_head_to_head_span_result, gated onsince/careerrather than folded
into the single-season branch, so that branch's own wording is untouched.
The answer names the actual seasons the games came from (_Span.during),
not only what was asked, and says "lead the all-time series" rather than
"won the series" for an unfinished, ongoing span - "won" reads as a settled
result, which a since-bounded or career tally is not. A stray player name
inteams(ahead_to_headquestion about "Embiid's record", which the
router resolves to his team) used to refuse for the wrong reason (span
unhonored); it now reaches the real one - a team-shaped name that fails to
resolve says so, and a real team pair answers.Golden-verified:
INTENTS=team_record,head_to_headover the recorded
corpus plusconstructed_cases_team_cells.jsonl(~/association-research/algebra-spike/step3/,
1989-1993 postseasons and since-bounded spans included) - 58 cases, 41
identical before/after and 17 changed, every one an unhandled-slot refusal
turning into an answer (or, for one recorded case naming a conference
situation alongsidespan, a refusal narrowing to its one real remaining
cause). One-token perturbation (disabling_record_narrowed'ssince
branch) caught by the harness in 6 of the 8since-bearingteam_record
cases - the other two ask "since 1989", the postseason's own floor, so
removing the clamp admits no extra games and the two runs agree by
coincidence, not by the guard failing to matter;head_to_head's own
since/spanpath does not call_record_narrowedat all, and none of
its cases moved, confirming the two templates'sincehandling are
genuinely separate code. Independently verified against the warehouse: the
Celtics' 2022-on regular-season record summed straight fromreal_games
(289-121) matchesteam_record's own count exactly.tests/query/test_templates.pygains a dedicatedteam_cells_confixture
(three regular seasons and two playoff series, one from 1989) and one
warehouse-verified test per newly-honored cell:sincealone,sincewith
an opponent,sincewith a venue,sincewith no games,game_nwithin
one named postseason,game_ncombined withsinceacross the 1989 floor,
game_nrefusing a regular season, the three conflicts (since+season,
since+career,since+month split) forteam_record;since,span
"career",sincereaching the 1989 postseason floor,sincewith no
meetings,sincewith a venue, and the two conflicts (since+date,
since+career) forhead_to_head. -
The condition skeletons discard the relation's window; the team relation
reads a barelimitthe way the player one does. With C5 reading a bare
limitas "the newest N" inscoped_games, the router's fillerlimit: 1
on "76ers record when Maxey scores 20+" cut a 63-game record to one game -
caught by the step 3 golden set (3 of 323 recorded cases moved), not the
suite.common.whole_spansays once whatRELATION_SCOPING_EXCLUDED
already declares forplayer_splits,record_whenandstreak- a split,
a record or a run is read over every game in the span - and
condition_playerand the three team branches apply it.team_gamesnow
sets its window through the same_relation_windowrule, so a team
question that dropsorderthe way the router does for shots ("last 2
games" as a barelimit) is windowed rather than answered for the season;
closes the ISSUES.md entry that recorded the two relations disagreeing. -
Shots are read through the player-games relation, and the window belongs
to the relation now, not to any one template (step 3, C5).shot_chart
andshot_distance(templates/shots.py) used to narrow by hand:
_scoping_gamepicked one game fororderfromplayer_game_logwith its
ownORDER BY/LIMIT 1,_shot_distance_wherewrote
athlete_id = ? AND season = ? AND season_type = ?itself, and
shotchart.render_for_playertook a singleevent_id- soordermeant
"one game" rather than a window, andopponent,venue,date,
without,split,since,game_n,below/aboveall refused.Both templates now settle their player and span the way every other
relation template does -shot_distancethroughcommon.scoped_player;
shot_chartthrough its own_shot_chart_settle_player, which keeps
shotchart.resolve_chart_player's best-match handling (a chart of the
wrong Curry is obvious on sight) rather thanscoped_player's strict
refusal between candidates - and read their games through
common.scoped_games/player_games.games_subquery(namedto bind it),
the same relation reader every other per-game template on it uses.
HONORED_SCOPINGfor both is now the full relation set with no
exclusions - an opponent, a venue, a teammate's absence, a starter/bench
half, one game of a series, a line on a box-score column, one Eastern date
andsinceall narrow which games are drawn or averaged, where they used
to refuse.The window itself moved to the relation.
player_games.Narrowed.window
("the newest/oldest N of the narrowed games, cut after every other filter")
existed already but nothing set it;common.scoped_gamesnow sets it from
order/limit(common._relation_window), andplayer_games.games_subquery
now honors it (routed through the existing_windowedreader
aggregate_sql/grouped_sqlalready used) - a no-op for its other callers
(player_splits,streak,with_without,record_when), which never set
a window, so_windowedreturns the identical unwindowed read it always
built.Narrowed.filters()grows an opt-inwindowedparameter (default
False):.windowis now set from every caller's slots that carry
order/limit- includinggame_logandperiod_split, which already
say "last N games" their own way (rows_sql, unaffected, never reads
.window) - so including the phrase unconditionally would have said it
twice in theirs. Onlyshot_chart/shot_distancepasswindowed=True.The headline fix: "Create a shot chart for Steph Curry's last two games of
the regular season" used to draw the whole season (374/803 attempts)
becauselimitwas never read onceorderhad scoped to a single game; it
now draws the two games (10/22 attempts, "in the 2026 regular season over
his last 2 games"). The router's own traces for this exact question,
checked across four separate runs and three builds, never emit anorder
slot at all - onlylimit- socommon._relation_windowreads a bare
limitas "recent" (documented in its own docstring, since nothing on the
relation has any other use for one - it ranks nothing), which is what
actually reaches the real question rather than only theorder+limit
shape the finding was first written up as.Measured by golden comparison over
INTENTS=game_log,player_stat,period_split,shot_chart,shot_distance
(323 cases: the recorded StatMuse/yardstick-v2 corpus for all five intents,
plusconstructed_cases_c5_shots.jsonl):game_log(144),player_stat
(101) andperiod_split(40) are 100% identical before/after, confirming
the window's new home does not change any of their answers. Of the 15
shot_distance and 23 shot_chart cases, 12 real answer changes - each a slot
that used to refuse now answering, or "last two games" now drawing two
instead of the season; the rest tempdir-path noise. Confirmed by a
one-token perturbation of_windowed's sort direction (player_games.py),
caught by the harness in exactly the 7 cases that read a window through it
and nowhere else. Newtests/query/test_shot_scoping.pypins a
warehouse-verified fixture case for every newly-honored cell.Fixed in passing:
_career_shot_notewas gated onspan.careeralone, so
asince-bounded read ("since 2024") claimed to "cover his whole career on
record" against the player's real, unbounded range - now gated on
span.career and span.since is None.tests/query/test_templates.py'stest_scope_guard_refuses_what_the_question_text_narrowed_to
dropped itsshot_distancecase (it asserted the old refusal, which this
step replaces with the capability) andtest_scope_guard_lets_through_what_the_player_templates_now_honor
and a newtest_shot_distance_narrows_by_opponent_and_names_it_in_the_answer
gained it back the other way - confirming an opponent narrows the average
and is named in the answer, rather than only thatcheck_scopelets it
through. -
period_splithonorsgame_n, as it always claimed to. Its shot read
handedscoped_gamesa dict built on the spot - venue, without, split -
instead of the question's slots, so "game 1 of each series" and the whole
postseason answered identically (LeBron James, 2018 postseason: 207 points
over 22 games either way). The relation now gets the slots whole, and the
source gate that stops a template narrowing the relation by hand also
refuses a hand-built dict handed toscoped_games/condition_player-
watched to fail on the old line. Found by the C5 period agent and filed;
fixed at the merge. -
period_split's per-quarter shot read no longer needs a literal season
to run (step 3, C5, part 1)._period_split_rowssummedshot_chart
overseason = ?/season_type = ?, taken from the already-settled
span.season- which isNonefor a career, and bound as SQL that matches
nothing.span"career" was already declared honored
(HONORED_SCOPING["period_split"]), so this was reachable, not
theoretical: a career question for a player with games on record answered
"no games found", the same false-cause shapeAGENTS.mdwarns about
elsewhere. The shot-value CTEs now join to the relation's own selected
games byevent_idinstead, which needs no season parameter at all.
Pure refactor for every question already reachable through a single named
season (proved by a golden comparison over the recorded and constructed
corpora - see the next entry), and a genuine fix for the career case,
which the next entry's own reasoning then refuses to answer for a
different reason. -
period_splithonorsdate, and refusesspan/sincefor a reason
worth stating outright (step 3, C5, part 2). One calendar day narrows to
its own game the same way it does forgame_log, through
common.scoped_games, and the answer says the date - the season used for
the reconciliation caveat is read off that game rather than guessed from a
season slot the router usually defaults to "now".span"career" and
sinceare refused rather than answered, even though part 1's join fix
would answer them correctly now too:PERIOD_RECONCILIATION, the
per-quarter accuracy caveat this template exists to attach, is measured per
season, and a range mixes seasons of different reliability under one
caveat (or none), with the header still naming a single season regardless- measured,
since=2023pulled the correct 257 games back to 2023 while
still heading them "the 2026 regular season".period_leaderboard's own
no-player read of the relation was assessed and NOT ported - documented in
its docstring and filed inISSUES.md("period_leaderboardstays off
the player-games relation"):opponent/venue/datewould need a new
no-player narrowing step this module may not add tocommon.py, and
since/spanreopen the identical per-season reconciliation problem just
fixed above. The twostep 3, C3gates
(test_templates_on_the_relation_declare_no_scoping_of_their_own,
test_templates_on_the_relation_do_not_narrow_it_themselves) now cover
shot_chart/shot_distanceonce the parallel branch porting them has
merged, detected at run time (_c5_shots_ported) rather than needing a
second edit here.
- measured,
-
The team-games relation grows a window and one game of a series (step 3,
C4b, part 1).TeamNarrowed(query/team_games.py) gets the two cells
C4 left off, mirroringplayer_games.Narrowed:window(the newest or
oldest N of the narrowed games, cut after every other filter -_windowed,
androws_sql/aggregate_sql/games_subquerynow read through it; unset,
every one of them emits the exact SQL string it did before, so this is a
pure refactor - confirmed by the full offline suite passing unmoved) and
series_game(one game of each playoff series, numbered overreal_games
by_TEAM_SERIES_GAMES, the team relation's own copy of the player
relation's_SERIES_GAMES).TeamNarrowed.filtersgetsopponent=/date=
toggles for a caller that already names either its own way - infrastructure
only here; nothing yet sets either new cell, so no template's answer
changed there. -
Scoping is declared once on the team-games relation (step 3, C4b, part
2).team_record,team_leaderboard,head_to_headand
team_quarter_pointseach still declared their ownHONORED_SCOPING
frozenset even after C4 put them (andplayer_splits/record_when/
streak's team branches) onquery/team_games.py's relation - the same
drift C3 removed for the player-games relation.TEAM_RELATION_SCOPING
(opponent,venue,date,since,span,order,game_n) is that
declaration for the team relation,TEAM_RELATION_SCOPING_EXCLUDEDthe
per-template reasoned exceptions, and_team_relation_scoping(intent, *extra)builds a template's set the way_relation_scopingalready does
for the player side.team_record's andhead_to_head's new declarations
are byte-identical to their old hand-written frozensets ({"venue", "opponent", "span", "situation", "split"}and{"opponent", "venue", "date"}) - confirmed by_team_relation_scoping(...)equality, so this
part is a pure rename of the mechanism, not a behavior change. Two
source-reading gates in the newtests/query/test_team_relation_scoping.py
(kept apart fromtests/query/test_templates.py, another agent's file
this round) mirror C3's: every declaration must equal
_team_relation_scoping's own output, andhead_to_head/
team_quarter_points(the two fully oncommon.team_games) may not write
the relation's owntg.opponent_id = ?/tg.side = ?/
tg.eastern_date = ?clauses themselves.team_recordand
team_leaderboardstill narrow an opponent and a postseason venue split
by hand (_games_record_games,_venue_records) - real, pre-existing
exemptions, each watched by its own test that asserts the hand-written
clause is STILL there, so a future port onto the shared step fails the
assertion instead of silently leaving a stale exemption in place.Part 1's
TeamNarrowed.window/.series_gameare now wired in:
common.team_gamesreadsgame_nandorder/limitoffslotsthe
same way it already readvenue, andcommon.scoped_teamreadssince
the wayscoped_playeralready does - the team relation's own
since-bounded career was ALWAYS available through_span_of's existing
sincebranch (used since 4.3.0's_player_splits_team); what was
missing was every OTHER team branch reading it and_team_span_label
saying so (below).Three templates gain real capability from this, each closing part of an
ISSUES.md gap:team_quarter_pointsnow settles its team and span through
scoped_teamand reads every game throughcommon.team_games, instead
of a hand-written join overteam_box_stats/real_games. It honors
opponent(unchanged),venue,date,since,span("career") and
game_nfor the first time, andorder/limitas a window cut before
the linescores are summed - closing "show sixers first quarter scoring
for their last 10 games" and "trailblazers stats last 10 games 3 point
average 1st quarter" (ISSUES.md, live yardstick-v2 failures), both of
which refused for want oforder/limit. The answer now says every
narrowing that applied (TeamNarrowed.filters()), including the window
("over their last 10 games"). The relation reads its games from
real_gamesrather than joiningteam_box_stats; measured against the
2026-09-22 warehouse this changes no coverage -team_box_stats.home_away
and.opponent_team_id, the two columns the old join actually needed,
are populated for every 2013-2018 Chicago/New Orleans game whose OTHER
box-score columns are empty (94/94, 87/87, ... rows; AGENTS.md, "Whole
team-seasons of box scores are empty"), so the old inner join already
included them and this port changes nothing there. The router's own
ORDER_INTENTS(query/router.py) gains this intent too - code-side
post-processing only, not aROUTER_PROMPT/ROUTER_SCHEMAchange, so no
other question's routing can have moved.team_leaderboardhonorssincefor the two record metrics
(record,losses) - "nba team with least playoff wins since 2022"
(ISSUES.md) - tallied by_team_leaderboard_since_records, a grouped
query over the relation keyed by the team's CURRENT display name (a
multi-season total has no one season to keyseason_name_sqloff).
Every other metric still refusessinceby name: a season line
(team_season_stats) has no way to sum across a span of seasons yet.record_when's team branch andstreak's team and league-wide
win/loss branches honorsince- closing the last open part of
ISSUES.md's "record_when's team branch and streak's team/league branches
still refuse ..." entry for that cell (rewritten below to what remains).
record_when's team branch also honorsgame_n(a threshold record
narrowed to one game of each series is a real, answerable shape);
streak's team and league branches do not, and refuse it by name
(_condition_needs_player_refusal("streak", slots, "game_n")) - the
gamesgame_nnumbers are not consecutive to each other, so a run over
them would silently answer a run over a scattered subset rather than the
real games in between.without,split,season_n,belowand
aboveremain refused for both intents' team-only or league-wide shape;
_CONDITION_PLAYER_ONLY_CELLSdropssinceandgame_nnow that they
are real answers, and gains an*extraparameter so the two intents no
longer have to share exactly one refused set.
_team_span_label/_team_span_floor_note(templates/splits.py, shared
by_player_splits_team,_record_when_team_answerand_streak_team)
readspan.sincefor the first time - fixing a label bug that predates
this commit: a since-bounded team span is career-SHAPED (span.seasonis
None, same as a plain career), so_player_splits_team, which has called
_span_of(..., since=...)since C4, already rendered a since-bounded
answer with the same "every regular season on record (1994 onward)" label
a plain career gets, with nothing saying the question had named a starting
year at all; it now reads "since 2022 (2022-2026 regular seasons)" like
every other since-bounded answer in this project.Golden comparison for the parts that touch runtime code already in use
(common.team_games,scoped_team): the new cells are all additive
(sincedefaults to not being read unless a caller passes it, and
window/series_gamewere already proven additive in part 1), so every
existing call site's generated SQL is unchanged for a question that does
not set them - confirmed by the full offline suite (uv run pytest -q -n auto) passing unmoved.team_quarter_points's rewrite is not a pure port- its internal read changed shape entirely - so every pre-existing test
for it (tests/query/test_templates.py) is kept passing unchanged in
meaning, and new warehouse-shaped fixture tests cover each newly honored
cell (tests/query/test_team_templates.py,tests/query/test_conditions.py).
One necessary, minimal, additive fixture change:tests/query/test_templates.py's
tq_congainedneutral_site/venue_citycolumns on itsgamestable
(bothfalse/a real city on every row, changing no existing assertion) -
team_games's relation reads both unconditionally (cup-final detection,
the home/road split) and this fixture predates any reader of the relation
reaching it; every sibling fixture the relation is read against already
carries them. Perturbed and watched to fail: a token removed from
TEAM_RELATION_SCOPING_EXCLUDED's reason, an extra cell added to a
declaration, and a hand-writtentg.opponent_id = ?added inside
team_quarter_pointswere each caught by the new gates in
test_team_relation_scoping.py; the two "exemption" gates were confirmed
to fail when the narrowing they watch for is removed.
-
A one-game streak is "1 game", not "1 games".
streak's headline, for
a named subject and league-wide alike, pluralized by hand; a team's streak
narrowed to one opponent (new in C4 below) is the first common way to
produce a run of one, so it showed. -
A team's games are defined once (step 3, C4). "A team's games" was
defined three times, and two of them disagreed with the third:
team_metrics.TEAM_GAMES_SQLselects a postseason by the CALENDAR YEAR it
was played in and excludes the NBA Cup final from a regular-season record,
whichtemplates/games.py's own hand-written team-log query already got
right but restated separately, and whichconditions._team_games(read by
player_splits/streak's team branches and, until now, bywith_without)
still gets wrong - answering the wrong year for every postseason 1988-1993
and counting the Cup final as a regular-season game. The correct
definition now lives once, inquery/team_games.py
(TeamNarrowed, mirroringplayer_games.Narrowed, with the same four
readers -rows_sql,aggregate_sql,games_subquery,named- over
TEAM_GAMES_SQL's relation, moved here fromteam_metrics, which
re-exports it), and two shared steps ontemplates/common.py
(scoped_team,team_games) settle a team's name and span the way
scoped_player/scoped_gamesalready do for a player.
game_log's team half andhead_to_head(templates/games.py) and
team_record/team_leaderboard's ownTEAM_GAMES_SQLreads
(templates/teams.py) are ported onto it;games._TEAM_GAMES_SQLis
deleted. Proved a pure refactor by golden comparison over 550 cases
(~/association-research/algebra-spike/step3,INTENTSextended to the
six team-facing intents,constructed_cases.jsonlextended with 1988-1993
postseasons and 2023-2026 regular seasons) - identical before and after,
confirmed to fail under a one-token perturbation of the shared narrowing.
with_without's window read (conditions._with_without_games) is now on
the relation too:_with_without_team_gamesbuilds aTeamNarrowedover
every team a player's windows touch at once (list_contains, since a
traded player's windows can span more than one team) and reads it through
games_subquery/named, sinceconditions.pycannot importtemplates
(its own module docstring) to reachteam_games's_team_span_clause
directly -_with_without_team_span_clauseis a private copy of that one
clause instead. Measured on the 2026-09-22 warehouse, nowith_without
answer moves: every 1994+ row the two sources produce agrees exactly (the
old read already joinedreal_games, and no game in it lacks a
team_box_statsrow to join), and the postseason-label fault this
correction reaches is 1988-1993, whichwith_without's own coverage floor
(capped at 1994 byplayer_box_stats) already refuses - so nothing it can
ask reaches the seasons where the two definitions disagree. Confirmed by a
golden comparison over the same 482 cases (before1.json/after1.jsonin
the step's own directory) and a one-token perturbation of the new
list_containsclause, which the same harness caught (22with_without
cases moved).
player_splits,record_whenandstreak's TEAM branches (the callers
conditions._team_gameswas left for) are now on the relation too:
_player_splits_team,_record_when_team_answerand_streak_team
(templates/splits.py), plus the league-wide win-loss streak
(_streak_league_by_result), readcommon.team_games/TeamNarrowed
instead of the label-scoped_team_games. This port is proved a pure
refactor separately (golden comparison over the same 103 team-condition
cases, held at the OLD_team_gamesbehavior for it by two interim-only
functions,_team_misfiled_postseasonand_team_scope_interim_floor-
101 of 103 identical, the remaining two a formatting-only change:
player_splits' team-only opponent/venue phrase now reads
TeamNarrowed.filters()'s "vs the X"/"at home" rather than the old
hand-built "(vs the X)"/"(at home)", the same wording every other template
on either relation already uses). Withwith_without's
windows ported in the same step (above),conditions._team_gameshas
no callers left and is deleted; the last of the three
definitions is gone. -
A team's postseason before 1993-94 is read by the calendar year it was
played in, forplayer_splits,record_whenandstreak's team
branches, andstreak's league-wide win-loss run. The interim guards
from the refactor above (_team_misfiled_postseason,
_team_scope_interim_floor) are deleted, letting the team-games relation's
own calendar-year read (already correct forteam_record/team_leaderboard/
game_log/head_to_head) reach these four for the first time. Two
behavior changes, both DATA.md-documented ESPN faults these branches used
to get wrong or refuse over instead of answering:- A single named postseason before 1993-94 used to be refused outright
("The warehouse files playoff games from before 1993-94 under the year
the season began..."); it now answers with the real games played that
calendar year. Measured against the 2026-09-22 warehouse:streak(team="Los Angeles Lakers", kind="win", season=1989, season_type=3)used to refuse
and now answers "The Los Angeles Lakers' longest winning streak, 1989
postseason: 11 games, 1989-04-27 to 1989-05-28" (their real run to that
year's Finals);record_when(team="Chicago Bulls", stat="points", threshold=100, season=1991, season_type=3)now answers "12-0 at 100+,
3-2 under, 15-2 overall" over their real 17-game 1991 title run, where it
used to refuse. 20 constructed cases over the 1989-1992 postseasons
exercise this (constructed_cases.jsonl; DATA.md, "ESPN files every
season before 1993-94 under the year it STARTED"). - A team's CAREER-wide postseason span used to stop at 1994 even though
the team tables reach back to 1988-89.conditions._game_scopeforced
every span (not just a single named season) tomax(1994, floor),
because there was no calendar-year read available to trust for the
earlier years; the relation's own floor is 1989 (nba/coverage.py,
postseason_first_season) and this is the first team-condition template
to read it for a career. Measured:player_splits(team="Detroit Pistons", span="career", season_type=3)moves from "1996-2026
postseasons (174 games)" to "1989-2026 postseasons (231 games)" - the
Pistons' 1988-89 and 1989-90 championship runs (57 games) were simply
absent from a "whole career" answer before.
player_splits' team branch already had a workingopponent/venue
narrowing (4.3.0);record_when's andstreak's team branches gain it
here, reading it the wayteam_recorddoes and saying the narrowing in
the heading viaTeamNarrowed.filters()- closing the ISSUES.md entry
("record_when's team branch and streak's team/league branches refuse the
relation's cells rather than reading them") for these two cells;since,
season_n,without,split,game_n,belowandaboveare still
refused for a team-only or league-wide question, and that entry is
rewritten to say so.streak's league-wide branch (nobody named at all)
still refusesopponent/venueby name
(_streak_league_needs_named_subject) - a league-wide streak has no
single team's rival or home/road split to read. Measured against the
2026-09-22 warehouse:record_when(team="Boston Celtics", stat="points", threshold=110, opponent="New York Knicks", season=2026, season_type=2)
answers "1-0 at 110+, 0-3 under, 1-3 overall" over their 4 meetings (used
to refuse);streak(team="Boston Celtics", kind="win", venue="home", season=2026, season_type=2)answers a 6-game home winning streak, still
open (used to refuse).Golden comparison over the same corpus this step's refactor commit used:
81 of 103 cases identical (unaffected by any of this), 20 changed - the
constructed 1988-1993-postseason and opponent/venue cases this commit adds
meaning to; none of the pre-existing 81 moved._condition_team_no_games
(added in the refactor commit, unchanged here) is what makes a team with
real games in a span but none matching a named opponent/venue get that
sentence now that the narrowing can apply at all, rather than the old
single-tier "no games in this span" a bare zero-row result used to print
regardless of which fact was actually missing. - A single named postseason before 1993-94 used to be refused outright
-
The scoping matrix cannot grow back (step 3, C3). Two tests read the
templates' source: the six on the player-games relation must declare
exactlyRELATION_SCOPINGless a reasoned exclusion, and none of them, nor
any private step it reaches, may narrow the relation itself - no call to
the narrowing function, no hand-written clause on an opponent, venue,
starter or date column. Both watched to fail.scoped_gamestakes an
already-resolved opponent as well as a name, which removed the two places a
template applied the opponent's id by hand around it. -
Scoping is declared once, on the player-games relation (step 3, C2).
The six templates that settle a player and narrow his games through the
shared steps -game_log,player_stat,period_split,player_splits,
record_when,streak- declared what they honored one list at a time,
and the lists had drifted: twelve slots on one, one on another, on the same
relation.RELATION_SCOPINGis the one declaration, and
RELATION_SCOPING_EXCLUDEDnames the few cells a template refuses with the
reason ("one game is not a run").player_statnow honorsdate- one
game, read as that night's line ("had 33 points, 3 rebounds and 6 assists on
2026-03-01"), found by its Eastern date and refused, naming the date, when
he had no game. The relation gained the window C0 ruled as its one
skeleton-specific rule:Narrowed.windowis the newest or oldest N of the
narrowed games, cut after every row filter and before an aggregate, so a
future "his average over his last 5 vs Boston" averages the five Boston
games; no template sets it yet (player_statstill hands "last N" to
game_log, by decision).
record_whenandstreaknow claim every cellRELATION_SCOPINGdeclares
for the six templates on the player-games relation (opponent,venue,
without, a named half of the starter/benchsplit,game_n,since,
season_n,below,above) rather than the barespanthey honored
before, and both read every one of them through the shared steps they
already settled through (condition_player), the wayHONORED_SCOPING
claims them.condition_playertakes ameasuresparameter (default none,
existing callers unchanged) so abelow/aboveline built with
measure_filtersin the template body narrows the pool the same way an
opponent or a venue does -record_when's own threshold stays the split,
and a measure line narrows what it is read over ("20+ points AND 5+
assists"). Both answers now say what they narrowed to via
Narrowed.filters()in the title, andsince/season_nget the same
phrasinggame_log/player_statgive them ("since 2022 (...)", "in his
18th season (...)") through a new_condition_span_label, since
record_whenandstreakbuild their heading off the relation's own
_Scoperather than the_Spancondition_playerresolves internally.
streak's run is read over the narrowed games, so "longest run of 20+
point games vs Boston" is a run over his Boston games only, and the answer
says so. The team branch ofrecord_whenand the team/league branches of
streaksettle no player, so none of these cells reach them yet;
_condition_needs_player_refusalrefuses one by name there rather than
silently answering as though it had been applied (ISSUES.md has the
follow-up).
player_splitsandperiod_splitnow claim every slot
templates.common.RELATION_SCOPINGdeclares (minus a per-template
exclusion inRELATION_SCOPING_EXCLUDED, each with its own reason) rather
than a hand-picked subset, and each newly claimed slot now actually narrows
the answer rather than being silently dropped:player_splitsused to handcondition_playera copy ofslotswith
"without": Noneand"split": None- so "Embiid splits without
Harden" and "Tatum's numbers as a starter" reachedcheck_scope's
declaration and nothing else.without, one game of a playoff series
(game_n), a range or ordinal season (since/season_n) and a line on
a box-score column (below/above) now reach the relation, and a named
half of the starter/bench split (splitas"starter"/"bench") now
narrows the games while the CATEGORY shown stays the one it always was -
both groups side by side, folded back from the half the question named
(previously this always raised, since "starter"/"bench" are not in
SPLIT_KINDS). The heading now names every narrowing through
Narrowed.filters(), the same phrase every other template on the
relation renders, rather than avenue/opponent-only phrase this
template composed for itself ("at home", not "(at home)"; the opponent
before the venue, not after).period_splitnow honorsbelow/abovethe same way: a line on a
box-score column narrows which of the player's games are summed for the
quarter or half, throughcommon.measure_filtersandscoped_games
rather than a hard-codedmeasures=[].common.condition_playertakes an optionalmeasuresparameter (default
None->[], so its other two callers are unchanged) to reach
scoped_games's own;common._condition_scopenow refusessince
alongside a namedseasoninstead of silently preferringsinceand
dropping the year, the same pairing_span_ofalready refused for
game_logandplayer_stat.
-
The steps that settle a player and his games are written once (step 3,
C1).game_logandplayer_stateach wrote out the same sequence - settle
the span, resolve the name against it, settle an ordinal season once he is
known, then narrow his games by opponent, venue, an absent teammate, a
starter/bench half, a playoff-series game, lines on box-score columns and a
date - and a fix to one had to be found and repeated in the other.
templates.common.scoped_playerandscoped_gamesare those steps, read
from the slots in one place, so a narrowing the relation learns reaches every
template built on them.player_splitsandrecord_whenfollow through
condition_player, which settles the player the same way and hands his
games to those templates as the relation renders them
(player_games.games_subquery) - the second reader of a player's games,
conditions._player_games, now has one caller fewer per port.streak
follows through the same step (player_games.namedrenders the relation
with named parameters, for a reader that nests the same subquery twice). No answer changes: proved by a golden comparison of
460 recorded and constructed slot sets across the ten templates that read a
player's games (answer text, data and refusals identical), with the
comparison watched to fail when one slot was dropped from the shared
function (14 cases moved).period_splitis now on the same two steps -
_period_split_rowsreadsscoped_gamesfor its narrowing rather than its
own call to_narrow_player_games, andperiod_splititself reads
scoped_playerfor the name and the "current or named" season it already
read one way - with the opponent still resolved eagerly through
_optional_teambeforehand, since the answer needs its name whether or not
any games end up narrowed to it. Proved by the same golden comparison
(460/460 identical), watched to fail whenwithoutwas dropped from the
narrowing (6 cases moved).
function (14 cases moved).
threshold_countandsingle_game_highwere assessed against the same two
functions and left on their own machinery, documented in place rather than
forced: both have a league-wide mode with no player named, which
scoped_player's name resolution cannot express, and both read from
league()- an everyone-at-once relation a singleathlete_idfilter
narrows to one man - rather thanscoped_games, which always narrows one
already-resolved player's rows and offers filters (opponent, venue, an
absent teammate, a split, a series game, a date) neither template honors.
Their own season/career check (_career_span) also stays: it raises
different wording thanscoped_player's_span_ofon a malformed span, and
onseason == 0where_span_ofwould not, so swapping it in would be an
answer change even though today's 460-case corpus does not happen to reach
either path (confirmed by perturbing the messages and re-running the golden
comparison, which came back identical). -
The router no longer invents a
seasonor adatethe question never
states. A bareseasoninteger ordatefrom the model used to be
trusted on its own whenever the question named no year or day - measured
live, "show me stats for sixers when maxey scored 20+ points" arrived with
season=2023 (nothing in the text but "20+") and answered a real player's
real average for a season nobody asked about, and "fingerprint maxey vs
jaylen brown 2026" arrived with date='2026-01-01' and was refused ("not yet
for a particular date") for a cause the question never gave. Both are now
kept only when the question itself supports them -_validate_seasonstill
reads a year or "last season" from the text first, and now drops a leftover
model guess rather than trusting it;_route_calendar_slotsdrops a
model-supplieddateunless a real calendar day is in the question. A year
or day the text does name (season_ref, "last season", an ordinal season,
a stated year, a stated calendar day) is unaffected. Fixes #95. -
"Last N games" with no season type named now reads both types. A "most
recent games" question that never says "playoffs" or "regular season" used
to default to the regular season alone, silently - "Show me the Knicks last
5 games" listed games through April even after the team's season carried on
into the postseason, and "what did Nikola Jokic do in his last 5 games?"
left out the playoff games it should have named.game_lognow reads both
season types for that shape and merges them by date, and says in the
heading what it found: "last 5 games of the 2026 postseason" where every
kept game is one type, "last 6 games (1 regular season and 5 postseason)"
where they are not. "Last 5 regular season games" and "last 5 playoff
games" are unchanged - saying the type outright is how the default is
corrected. Read from the question inrouter._route_game_log_recent_span,
the same waysideandcoachare, so no other question's routing moves. -
A name the question leaves open means whoever still plays, and the answer
says so. A name several players share used to ask "which one?" whenever
more than one of them had a row in the seasons the answer could read - so
"show maxey's games against boston in the past two seasons" asked about
Marlon Maxey, who retired in 1994. With no season asked about, the one
namesake who played the last season of the span is now the answer, and two
who both played ("brown", "curry") are asked about exactly as before. A name
given in full yields the same way where its owner has nothing in the seasons
asked about and exactly one namesake does: "Jabari Smith" for 2026 is Jabari
Smith Jr., where it answered "no 2026 games" about his father. This is a
default, so it is said: the answer ends "('maxey' was read as Tyrese Maxey,
the only match who played in 2025-26. Marlon Maxey also matches - use the
full name, or name a season he played, to ask about him.)", and the same
sentences are inAnswer.data["name_readings"]. Measured on the warehouse,
124 of the 391 surnames two or more players share stop asking, and in 67 of
them a retired namesake has more games on record than the active player
("wade" is Dean Wade, "pippen" is Scotty Pippen Jr.) - which is why the
sentence, with the wording that reaches the other player, is part of the
change rather than a nicety. New:entities.collect_name_readings. -
A router that could not be asked says so, instead of blaming the
question.route()collapsed three different failures into one: ollama
unreachable, ollama unable to serve the router model, and the model replying
with something unparsable all returned None, and the caller reported "the
router returned no usable classification". Only the third is that. Reported
from a laptop where the router model was not pulled: every question came
back with that sentence, which reads as a fault in the question and points
nowhere near ollama - and under--disable-fallthroughit is the entire
error. The first two now raiseRouterUnavailable, naming the model and the
server ("ollama could not serve the router model 'qwen2.5:3b': model not
found"). The rule that a router failure costs a round trip and never an
answer is unchanged:Agent._ask_innercatches it and falls through exactly
as before, andscripts/check_routing.pysays it once and stops rather than
reporting every case as a routing failure. -
The web page recalls earlier questions, and every question can be copied.
ArrowUp steps back through the questions asked on the page and ArrowDown
forward again, the way a terminal does: it holds at the oldest rather than
wrapping, stepping past the newest restores the half-written question it
interrupted, and a repeat of the question just asked is not a second entry.
It is only history when the caret is on the first line (ArrowDown, the
last), so a question typed across two lines still moves the caret between
them; typing anything restarts recall from the newest, because someone who
browsed to the oldest entry and then typed something fresh otherwise found
ArrowUp doing nothing at all. Every question balloon now carries a copy
button for its own text, usingnavigator.clipboardwhere the page is a
secure context and falling back to a selection-based copy where it is not -
association webprints a loopback URL, which IS one, but the same server
answers on a LAN address over plain http, where the API is simply undefined.
A failed copy says so rather than looking like nothing happened. The
questions live in memory only, so a reload clears them, like the answers
above them (ISSUES.md #69). -
A player's record against one team is his games, not two franchises
meeting. "Embiid career record vs boston" and "Show Embiid's career record
against Boston" both routed tohead_to_head, which counts every
76ers-Celtics meeting including the ones he sat out - and both fell through
before reaching that wrong reading. They arrive two ways, with the player in
theteamslist and with him replaced by his own team and named only in the
question, andentities.player_record_against_a_teamturns both into
with_without, whose split is exactly "the games he played against the ones
he missed". The router cannot make that call itself: whether a name in
teamsis a player or a franchise is a fact about the warehouse, not about
the words.with_withoutnow honorsopponentto go with it, narrowing
both rows of the split together and naming the opponent in the title, since
a record over one opponent's games headed as though it covered every game is
the silent narrowing that module exists to stop. Measured: 76ers 13-15 in
the 28 regular-season games Embiid played against Boston, 3-8 in the 11 he
missed. The handler is re-resolved with the intent: it shipped once without
that, so the trace readhead_to_head -> with_withoutand head_to_head ran
anyway and refused for wanting two team names, which every offline replay
passed because the replay script resolves the handler after the rewrite and
the pipeline resolved it before (#163). -
One definition of what a question calls a box-score column.
route()
reads the stat beside a threshold ("20+ points") andtemplates/common.py
reads the same phrases onto columns ("under 14 fta"), and each kept its own
copy of the vocabulary, becauserouter.pyimports nothing from
templateson purpose. Nothing checked that they agreed, and
check_duplicate_names.pycould not - the two names differ. Both now read
association.query.measures.MEASURE_WORDS, a leaf module with no imports of
its own; the router names only which SPELLINGS its threshold grammar
accepts, and takes what each means from there, so a spelling dropped from
the shared map raises at import rather than silently narrowing the grammar.
The rebuilt pattern was re-proved identical over the same 317 strings (#164). -
A subject keeps the first name the question gave it. The grammar that
reads a dropped subject back out of a question captured ONE word, so a
possessive gave a bare surname: measured over the 261-question corpus,
"kobe bryant's stats vs rockets" namedbryant- Bryant Reeves, Bryant
Stith, Carter Bryant and Elijah Bryant, and not Kobe - and "Jaden
mcdaniel's" and "steve adam's" did the same. It captures an optional
leading word now, as the count grammars already did, gated on the same
stopword list so "most points curry scored" still readscurryand never
points curry. A bare "had"/"has" before a number is a subject position
too ("sixers record when maxey had 10+ rebounds" named nobody, while the
same question with "scored" answered), and the stopword list gained the
function words that can sit before a name, which is what kept "when maxey"
from being read as one. Net over the corpus: three full names recovered and
four pieces of junk dropped, "game score nba leader" - which named a player
called "game" - among them. No corpus question's routing or answer changes
(#165). -
A ranking asked for in a unit nothing is stored in is refused, by name.
"who were the top 10 in defensive netpoints / 90" fell through to the agent:
ratewas in no template'sHONORED_SCOPING, socheck_scoperaised - a
sentence that reads as a refusal in the trace and is not one, since the
question then reached an agent with no per-90 anything to read and nothing
to stop it filling the silence.leaderboardnow declaresrateand
answers it: a season total where the metric has one, and otherwise a
refusal naming the metric's real forms ("No leaderboard ranks netpoints
defense per 90 minutes - the warehouse stores it only per game or per 100
possessions"), which is computed from the metric rather than listed
generically so it cannot offer a form that does not exist. The same
omission had madeleaderboard's existingrate == "total"branch
unreachable through the pipeline, so "most points this season" as a season
total is answerable for the first time (#152). -
A record "when X had 20+ points" reads its own threshold. Measured live
in the 2026-09-20 web session: "what was the sixers record this season when
tyrese maxey had 20+ points?" came back withstat='wins'- "record" is
what the model had to file understat, which is required and whose enum
has no won-lost record, so the nearest value it knew won - and the template
refused with "record_when needs a known stat and a positive threshold, got
'wins'/20" about a question that states its stat plainly.route()now
reads "20+ points" as the one fact it is and sets both halves for
record_when, the question's own words beating the model's guess; a
question stating two thresholds is left alone to refuse, since there is no
second threshold slot to put one in. The threshold vocabulary and the
pattern that matches it are now one definition - the alternation is built
from_THRESHOLD_WORDS- and the rebuilt pattern was proved to match
identically over 317 strings.ROUTER_PROMPTandROUTER_SCHEMAare
untouched (#114). -
A record "when X scored N" keeps X. "what was the sixers record when
maxey scored 15+ points?" routed torecord_whenwith the team, the stat
and the threshold all correct and noplayerat all, so the template raised
"record_when needs a player", the question fell through, and the agent spent
583 seconds producing nothing. The name was already sitting in the grammar
_SUBJECT_OF_HIGHreads ("maxey scored");record_whenwas simply never
asked, so it joinssingle_game_highandthreshold_countin
_SUBJECT_RESTORED_INTENTS. Restoring can only help here, where it could
not for the other two: arecord_whenwith no player is not a league
question but an unanswerable one. The question now answers 36-29 over the 65
games Tyrese Maxey scored 15+, matching the figure measured by hand on
player_game_log. -
record_whenanswers a team's own threshold, with no player named at
all. "what was the celtics record when they scored 120 points" names no
player by any grammar, so the template raised "record_when needs a player" -
the wrong cause for a question that was never about a player - and fell
through to the agent.record_whennow reads ateamslot with no
playeras the TEAM's own threshold:pointsis read straight off
real_games' own score rather than ateam_box_statsrow, so it needs no
box score at all and is immune to the empty 2013-2018 Chicago/New Orleans
team boxes; every other whitelisted stat (rebounds,assists,steals,
blocks,threePointFieldGoalsMade,fieldGoalsMade,freeThrowsMade,
fouls) readsteam_box_statsand says how many of the team's games it
could not see there.turnoversreadstotalTurnovers, whichDATA.md
establishes as ESPN's correct figure in every era; onlyminutesis refused
by name, because a team has no minutes total. Measured on the 2026-09-20
warehouse: the Celtics were 27-0 scoring 120+ and 29-26 under it in the 2026
regular season.record_whenstays inPLAYER_REQUIRED_INTENTS: a question
naming exactly one player still restores him, which is what keeps "76ers
record with 20+ points from tyrese maxey" - a wording the subject grammar
does not fire on - from being answered as the 76ers' own scoring, while a
question naming nobody reaches the team branch either way (#144). -
A single-game shot chart or fingerprint names the game, not just its id.
A single-game shot chart's subtitle readf"game {event_id}"(e.g. "game
401705764"), and its answer named only the player and the made/attempted
split - neither said which game had been drawn, not the date, the opponent
or the result, thoughgamesholds all three. The single-game fingerprint
shared half the shape: it already named the date but not the opponent or
result. Both now read "2026-01-25 @ MIN, W 111-85" in the subtitle and the
answer - "vs" for a home game and "@" for a road one, as every other log
here writes it, via the newassociation.query.game_label.game_label- one
definition read by both renderers, so they cannot drift into two
descriptions of the same game. The filename keeps the bare event id; only
the reader-facing text changed. A game with no usable box-score row (a
placeholder event, an unposted score, or a warehouse without
player_game_log) falls back to the old bare-id/bare-date text rather than
raising or printing "None vs None" (#155). -
A fingerprint "vs" note stopped naming a player it holds as one it does
not. "show a fingerprint for maxey vs jaylen brown in 2026" answered a
single polygon for Tyrese Maxey plus "only one of them matches anybody in
the warehouse - check the spelling of the other" - false:playersholds
Jaylen Brown andnet_points_player_fingerprinthas his 2026 row.
entities.restore_dropped_playerscomparedplayers_named_in's COUNT
against the held slots' count, and "Maxey" alone names two players
(Tyrese and Marlon), soplayers_named_in's own strictness dropped it from
its list - leaving one name on each side of the comparison and reading as
"nothing to restore" even though the two single names were two different
people. It now compares by CONTENT: a held name with any trace in the
question (entities._grounded) is kept, and whatever the question names
beyond that is added rather than used to replace the whole list.
entities.compared_but_unmatchedis now a second, independent guard against
the same false claim - it takes a connection and resolves the leftover name
against the roster before saying anything, so a resolvable name gets "was
not included in this answer" and only a genuine non-match still gets the
spelling note (#143). -
A team written with its space left out is still the team.
"trailblazers stats last 10 games 3 point average 1st quarter" arrived with
team='Portland Trail Blazers'correctly routed and then lost it:
entities._team_groundedasks whether the question holds a WORD of the
name, and the single token "trailblazers" equals none of {portland, trail,
blazers}, so a team the question opens with was dropped as one it never
mentioned - a regression from178c21f, which made an ungrounded team go
even where no player was found. The nickname was only the visible half: six
of the thirty teams have a two-word city, and_team_namedresolved none of
the seven run-together spellings, so_team_after_versuslost "vs
goldenstate" and "vs newyork" outright as well._run_togethernow derives
the spellings of a name with a space taken out, and grounding,_team_named
andfind_teamsall read them. It stays a lookup and never a guess: the
letters must equal a whole run of the name's own words in its own order, so
"la" and "trailblaze" still name no team, and all 44 forms of the 30 teams
resolve to exactly the team they spell with no collisions. Measured over the
261-question corpus: no answer changes, and the one question this reaches
keeps its team instead of handing the agent a question with no team in it
(#158). -
A team's quarter is refused for a stat the linescore does not hold.
Restoring that team would have answered the same question's "3 point average
1st quarter" with the Blazers' first-quarter POINTS - the fluent wrong
answer this architecture exists to prevent - becauseteam_quarter_points
readshome_linescores/away_linescores, which hold one number per period
and nothing else. Astatthat does not resolve to points now raises,
naming the linescore as the limit, so these questions fall through as they
did before rather than being answered about something else (part of #161). -
"All playoff games" draws every postseason, not just one presented as
all of them. "show a shot chart for steph curry in all playoff games"
routed to a single defaulted season (2025) and drew it - 62 of 130 shots -
with nothing in the answer saying it was one postseason out of the ten
Curry has (2013-2019, 2022, 2023, 2025; 3,866 postseason shot rows on record, 3,494 of them carrying coordinates). None of
_SPAN_WORDS("career", "all-time", "ever", "in/of history") is in "all
playoff games", so_validate_spannow also reads "all"/"every
game(s)" (anchored so "all-star" cannot fire it) intospan
"career", andshot_chart/shot_distancenow honor it - drawing (or
averaging) every season of the season type asked for, 1,387/3,055 makes
across Curry's whole postseason, and saying so. A career whose own start
predates the 2002 shot floor (COVERAGE["shot_chart"]) says which seasons
are missing rather than reading as though nothing were on record at all,
or, where the whole career is before it (measured: Michael Jordan's
1985-1998 postseasons), why nothing can be drawn (#141). -
"past two seasons" reads as a season span, not a games count. The
router's ownlimitis where a relative season count landed instead:
"show tyrese maxey's games against boston in the past two seasons" arrived
withlimit=2andspan="career", and answered his last 2 games of his
CAREER where 7 were asked for (3 against Boston in 2025 and 4 in 2026,
measured onplayer_game_log).route()now reads "past/last N
seasons"/"...years" intosince(current season minus N plus one, no
untilneeded since nothing is played after "now"), and drops alimit
whose only count word belongs to that phrase rather than to a real number
of games ("last 5 games in the past two seasons" keeps itslimit) (#140). -
Both conditions of a two-condition count are answered.
ROUTER_SCHEMA
carries onethreshold, so a second condition survived only as afields
entry the template ignores: "who had the most 30+ point 10+ rebound games
this year?" answered the 30+ point leader (Luka Doncic, 44) where the pair
is Nikola Jokic with 20, and "How many 20+ point 5+ assist games did luka
have?" answered 441 against 397.route()now reads every "N+ " the
question states and, where it states more than one, carries them all as
lines on box-score columns - the filters the relation already composes - so
threshold_count,game_logandplayer_statkeep only the games that
clear every one. A single condition is still the model's ownthreshold,
untouched, and the "+" is required, so "top 10 rebound leaders" is a
ranking and not a condition (#139). -
A chart is not narrowed to one game the question never named. On
shot_chart,shot_distance,player_netpointsandfingerprintan
orderresolves to a single event id, wheregame_logonly sorts, so a
filler one costs a whole season: "a shot chart of steph curry's 2025 season
for 3 point shots" drew one game, 7 of 12.route()now keeps anorder
on those four only where the question names a game at one end of the span.
And "steph curry's last regular season game" came back as season 2025 - the
model read "last regular season" as the season before this one, and the
chart drew a game a year off; a single game named with no year and no
season words now means the current season, which is the default every
template already applies (#153). -
A team's half is answered, and "most points in a half" is one game.
"Detroit Pistons most points in a first half this season" and "least points
scored by the wizards in the first half this season" fell through: a team's
half had no template, because the model maps "first half" onto period 1 and
that is wrong for a team the same way it is for a player. The linescore
already holds both quarters, soteam_quarter_pointssums them - it takes
thehalfslot through the same_period_scopethe player side uses - and
a question asking for the most or the fewest gets that single game rather
than the season's average, with every game named when two tie. Measured:
the Pistons scored 81 in a first half against Indiana on 2026-04-12, their
most, and the Wizards 37 against Chicago on 2026-04-07, their fewest. -
A quarter or a half is narrowed by the relation, so it honors a
teammate's absence and an order.period_splitkept its own copy of the
played-game guard, so "scottie barnes stats 2nd half log without rj" was
refused for a slot no period template honored while the relation had been
answering exactly that narrowing for four other templates since the port.
Its rows come from_narrow_player_gamesnow:withoutcomposes there
(with the teammate tenure rule and the clarifying question that comes with
it), andorderpicks which end of the season the log's rows come from
rather than always the most recent. The answer names the absent teammates,
as every other narrowing here is named. Proved a pure refactor over 20
recorded period questions - every answer byte-identical, and the comparison
watched to fail on a one-token change. Measured: Barnes played 80 games in
2026 and RJ Barrett played 55 of them, so his second half without him
covers the 25 the template reports. -
A quarter or a half can be ranked, not just looked up. "who has the
highest average 1st quarter points this season?" and "knicks 1st quarter
scoring leaders playoffs" reachedotherand fell through to the agent:
the router sends a period question with no named player there, and there
was nothing else to send it to.period_leaderboardranks players by their
points in one quarter or half, reading the same sourceperiod_splitreads
the same way - the value of each made shot throughSHOT_VALUE_SQL, never
the play's prose - with that template's season accuracy gating unchanged.
The denominator is games PLAYED, so a scoreless quarter counts as the zero
it is; a per-game average needs the qualifier every other per-game ranking
here applies (20 games, five in the postseason) and the answer names it; a
named team narrows the ranking to that team's players rather than becoming
the subject, and the ranking words win over the team's own quarter template
so "knicks ... leaders" is not answered with the Knicks' first-quarter
total. Measured: Luka Doncic led the league at 12.0 first-quarter points
over 64 games in 2026. -
"Record when X and Y played" is answered, not refused four ways. All
four phrasings a web session asked - "when both Embiid and Paul George
played", "with Embiid and Paul George", "when Embiid and Paul George play",
"when Embiid with Paul George" - routed torecord_when, which divides a
season by a NUMBER a player reached and so refused every one for want of a
stat and a threshold. A thresholdlessrecord_whenthat names players who
played together iswith_without's question, which already divides by two
teammates at once, and it now goes there: the 76ers were 13-11 in the 24
games Embiid and George both played. Both sides of "when A with B" are
read, since taking only the far side answered about George alone. A
question that does name a threshold keeps its intent (#156). -
A fingerprint pair is no longer refused over the router's own mis-slot.
"sga vs tyrese maxey fingerprint" put Maxey inopponent; the pair was
correctly rebuilt intoplayersand thencheck_scoperefused the
leftover slot, so a question the system answers under other words
("compare sga and tyrese maxey fingerprint") had no answer at all. An
opponentthat names no team and names a player the slots already ask
about is dropped. One that names a player nobody asked about stays, so a
template that cannot honor it still refuses rather than widening to every
opponent (#157). -
The fall-through agent gives up on a clock, and says what it could not
answer. An iteration cap never bounded the wait, because the cost is per
model call: measured over 24 questions, every finished run spent nearly all
its wall time inside 1-4 calls, 14 of the 24 never finished at all, and
one ran past 17 minutes before being killed by hand.Agentnow takes
budget_seconds(default 120,--agent-budgetonqueryandweb, 0 to
remove the bound), checked before each call so the first always runs. When
it gives up - on the clock or on the tool-call cap - it now names why the
templates declined the question instead of saying "Gave up after too many
tool-call iterations", which told a reader nothing about their own
question (#129). -
A faster check to iterate against.
scripts/check_fast.shruns every
gate except the Sphinx build plus every test except the one markedslow,
in parallel: ~39s against ~106s for the full check, which itself drops from
~180s now thatpytest -n auto(pytest-xdist, new in thedevextra) runs
in CI and locally. The oneslowtest is the Eastern-date agreement check,
57s of the suite's 80s; it still runs in CI and in every full local run. -
A NetPoints rate asked for by any name is ranked as the rate. "top 10
in defensive netpoints / 100 possessions", "adjusted defensive netpoints",
"offensive netpoints per 100 possessions" and "adjusted netpoints" all
ranked season totals - a different list (Wembanyama, Holmgren, Queta by
total; Wembanyama, Hartenstein, Capela per 100).route()now reads
"adjusted", "per 100 possessions", "per possession" and "/ 100" and
switches a NetPoints metric to its per-100 variant; a metric with no such
form ("points per 100 possessions") and a per-90 rate get arateslot no
template honors, so the question is refused rather than ranked by the
wrong unit (#152). -
--disable-fallthroughonqueryandweb, for development. When no
template can answer a question, the command returns an error saying why the
fast path gave it up (no usable classification, an intent with no template,
or the template's own refusal) instead of handing the question to the
SQL-writing agent, which iterates for minutes at a time and rarely gets it
right - "what was the sixers record this season when maxey scored 20+
points?" spent twelve minutes of a pegged CPU on five wrong attempts. The
CLI prints the reason and exits non-zero;POST /api/askanswers 501 with
it; the stream sends it as itserrorevent.Agentandservetake a
fallthroughargument, andFallthroughDisabledlives inquery.answer
so the web layer can catch it without a model client. The changelog hook
now also requires an## Unreleasedheading wheneversrc/changes: the
entry below landed inside the 4.3.0 section the day after that release. -
A team the router put in a player's place stays where it belongs. Four
fluent wrong answers from one entity stage: "karl towns stats vs netslast 5
games" restored Towns but dropped the Nets, answering his last five games
against anybody; "magic vs nets last 10" named Magic Johnson by the one
word "magic" and gave him the Magic's log; "Jersmi grant last 5 games vs
the suns" and "stating centers vs phoenix suns log" answered the Suns' and
the Lakers' logs, teams the question names only as the opponent or not at
all. Now the opponent stays the opponent, a word that names a team the
question is about is not a player, an "X vs Y" with nobody named is X's
log against Y, and an invented team goes even when nobody is named - so
the template refuses for want of a subject rather than answering for the
wrong one. And a name typed with accents ("luka dončić") is folded to the
plain letters the warehouse spells it in before it is matched.
These notes are cut short: the full 4.4.0 section (129980 characters) is longer than a GitHub release allows. Read all of it in CHANGES.md.