Repository navigation
Releases: jeffknupp/association
Release list
association 5.0.0
- Breaking (packaging): the distribution is
association-pyon PyPI; the import package, theassociationcommand and the repository keep the bare name. PyPI refusesassociationas a project name though nobody holds it: no project row (its per-project RSS route answers 404, as for a made-up name), no release ever, no name confusable with it under PyPI's rule over its full index of 909,700 projects, and no standard-library module - of the four causes PyPI's help page lists, that leaves an administrator's prohibition. Sopip install association-py(and'association-py[web]') is the install command from this release on;association.__version__anddocs/conf.pyread the installed metadata under the new name;association web --helpand the missing-extra hint saypip install 'association-py[web]'; the wheel isassociation_py-X.Y.Z-py3-none-any.whl. The releases before this one carry the old name in their metadata and stay on GitHub only.scripts/bump_version.pymoves a tag-pinnedassociation[web] @ git+...install line to the new name with its pin (EXTRAS_PIN), since pip and uv refuse aname @ urlinstall whose metadata name differs; one test. Trusted Publishing is unchanged: the publisher on PyPI is registered forassociation-py(AGENTS.md, "Releasing"). - Breaking (Python API): the answer side chooses a reader by the planned point's shape and a sayer by the body, never by an intent (the Phase 2 review's cleanup (b)6, "the Result is typed", third commit).
compose.answerpicked a reader byintent ==(72 mentions of "intent" incompose/__init__.py), passed every readerSTATED_SCOPING[intent](23 sites), andsay()was a decision list over the span's source, the relation, a facts key and the body. Now the planner settles what a point is read and said as -compose.plan.PointShape(relation, shape, by)(the target types' relation, shape andby;player_periodsandteam_periodsname a quarter's read), byplan.shape_of(intent, query), on the newPlanned.shape- andcompose.answerlooks the reader up in one table by that key alone (compose._ROUTES; a route markedonlyis the point's only answer and declines rather than step aside).STATED_SCOPINGis keyed byPointShape;plan.SHAPE_WORDSnames each shape's retired words, andplan.words_stated(name)is the planner's lookup while it still reads intents.say()chooses from the headline body's type and the one field that says what it is (Scalar.how,Grouped.by,Rows.by,Chart.kind) through three tables;Span.sourcechooses nothing (a team's record tells a tally from a standings line by its detail part). The intent is read by the planner only where the query alone cannot tell two retired templates' words apart (a log from a quarter's, the league's ranking from a quarter's, a count's shapes, a team's own total tried on its season line first;# Phase 3: needs the Reading's own shape and by), and asPlanned.floor, the coverage entry the point's season is checked against (coverage.SOURCESis still keyed by the retired words;# Phase 3: needs the relation's declared tables). "intent" incompose/__init__.py: 72 -> 6 mentions, all in prose; inquery/: 671 -> 635 (the planner gained what the answer side lost). Identical on 628 of 628 recorded questions, 2,082 of 2,082 feed answers and 2,710 of 2,710 readings, text and remarks; 2,556 of 2,556 unit-test calls identical in every outcome, and 981 of them differ in one argument by construction - theplannedvaluecompose.answeris handed carries the two new fields (shape,floor), checked by stripping them: 0 differ. Tests: 0 deleted; the scoping tests readwords_stated(name)andSHAPE_WORDS(9 files, by name), an empty log test gives its body; 1 added (the answer side names no intent - watched to fail). - Breaking (Python API):
Result.factsis typed per shape, and the cells a read applied are typed values on the narrowing and the span (the Phase 2 review's cleanup (b)5, "the Result is typed", second commit).Result.factswas aMapping[str, Any]written at 51 constructions with about 80 distinct keys and read by the sayer at about 120 sites (inventoried at run time over the 628 recorded and 2,082 feed questions,~/association-research/stages/facts_census.py, and statically,facts_inventory.py). Each key is now one of three things. A cell the read applied is a typed value: the newresult.Cellunion (Period,OnDate,Line,Role,GameOfSeries,Calendar,Companions,Met,ShotValue- the minimal set the keys needed, declared by this step) onNarrowing.cells, read by type (Narrowing.cell(kind),lines(),period), andSpan.since/until/ordinal; a split's or a run's parameter is its body's (Grouped.of: the line a record splits by, the teammates a with/without split divides by;Runs.line/won: what a streak held along). A fact the sentence needs is a field of one frozen record per shape (result.Facts:CountFacts,LogFacts,LineFacts,SplitsFacts,PeriodFacts,PeriodRankingFacts,TeamPeriodFacts,RankingFacts,RecordFacts,MatchupFacts,MeetingsFacts,TeamRecordFacts,TeamStatFacts,TeamRankingFacts,OutlookFacts,NetPointsFacts,ChartFacts), read by the sayer through one accessor that raises on a record of the wrong shape. A key that only chose the sayer is gone: a period split's"period" in factsis itsPeriodcell. Moved off the bag besides: theaboutpage values of a player's line and splits (now read off the cells), a ranking's season and a fingerprint's (the span's), a period log's order and limit and a period's window (Result.window), a shot distance's one game (anOnDatecell and the window it ended), the outlook's and the NetPoints season's "nothing found" (an empty body of the shape, where a Result with no parts was).Narrowing.periodis thePeriodcell (a label and its numbers, where it was a string);player_games.Narrowed.lineskeeps each line as values beside its words.Scalar.hownames every scalar's reduction (season,per_shot,per_100,total,projection,rankedbesideper_game,count,record). Identical on all four populations, text and remarks: 628 of 628 recorded questions, 2,082 of 2,082 feed answers, 2,556 of 2,556 unit-test calls, 2,710 of 2,710 readings. Tests: 0 deleted; 10 hand-built Results intest_say.pyandtest_templates.pyrewritten to the typed facts and cells; 1 added. - Breaking (Python API): a refusal found at RUN is typed -
result.Refusal(kind, facts)andresult.Clarify- and its sentence is the sayer's, one phrase per cause (the Phase 2 review's cleanup (b)1, "the Result is typed", first commit). Until now a reader that found nothing to say returned a wordedReply(36 constructions incompose/, 6 in the relation modules), aResultwith no parts andfacts["missing"](compose/team_stats.py, whichsaybranched on), or aResult.emptysentence; now every reader returnsResult | Refusal | Clarifyand nothing outsidecompose/say.pywrites a refusal's words.Refusal.kindis closed: a READ's causes are the newresult.RUN_CAUSES(50 kinds: a season under a floor, no games in the span, none matching the narrowing, a teammate never seen, a team short of games for points allowed, why no fingerprint can be drawn, ...), a READING's stayreading.CAUSES(17) - two disjoint sets, held by a test, because a reading's cause is decided from the words before any table is read and is pinned by the readings population, while a read's is a fact found in the warehouse; oneRefusaltype carries either, andsay.refusal_phrase(kind, facts)is the one table that words both (the planner's sentences moved there fromcompose.plan, whoserefusal_resultnow saysrefusal_of(cause)through it). A Refusal also says what the page shows beside the sentence (shown) and under which keys (under), since the readers' pages differed and the answers' data is held identical.Clarify(asked, candidates, kind, active, why)is the question back - several on record by a name, or a name nothing matched with near spellings. Moved or reshaped on the public surface:entities.clarificationandentities.suggestionarecompose.say.clarification/suggestion;entities.no_matchisentities.unmatched, returning aClarifyor aname_unmatchedRefusal;entities.clarify/resolved_player/resolved_team/optional_teamreturn aClarifywhere they returned a wordedReply;player_relation.no_narrowed_gamesandno_games,team_relation.condition_team_no_gamesreturn aRefusal;Result.emptyis aRefusal;compose.core.Refusedcarries the typed outcome;fingerprint.FingerprintUnavailablecarries a cause and its facts (kind,facts) rather than a sentence;coverage.floor_refusal/coverage_refusalarecheck_coveragetyped;say.say_period_refusal,say_period_unread,say_team_period_*,say_period_rank_*,say_conference_refusalandrankings.leaderboard_shot_distance_refusalare phrase-table entries. A combined record's half that was a refusal no longer has its words searched for wins and a "Note:" tail (none of the refusals a half can come back with carried either). Of the 614 answered recorded questions, the ones worded outside the sayer went 35 -> 7 (~/association-research/stages/typed_census.py); the 7 arerefusals.unanswerable/by_question, which read the question's own text in the answering loop and move with Phase 3's first item. Identical on all four populations, text and remarks: 628 of 628 recorded questions, 2,082 of 2,082 feed answers, 2,556 of 2,556 unit-test calls, 2,710 of 2,710 readings. Tests: 0 deleted; 13 rewritten to say the typed outcome (test_fingerprint.py9,test_entities.py5 via_no_match, one each intest_say.py,test_compose.py,test_templates.py); 2 added. - **A measure that reads no rebuilt line carries its own games...
association 4.4.0
- The subject reading's list of ordinary words ships with the package (
query/words.txt.gz, from Debian'swamerican2020.12.07, with its SCOWL notice inwords.COPYRIGHT) instead of being read from/usr/share/dict/words. On a machine without that file, "Best record from 2010-11 to 2018-19" was read as a question about Travis Best: the fallback kept only the month abbreviations, so any word that is also a whole player name ("best", "head", "pointer") became a player. GitHub's runner is such a machine, which is why CI failed from7b8b7a2on while the same tests passed locally. The shipped list holds exactly the 64,005 words the system file gave, so answers on a machine that has the file do not change. A wheel missing it now raises instead of silently answering differently. - The subject reading writes the first field:
subject.apply_subjectputs the players the question was read to be about intoplayer/players- a router name the question never held replaced by the question's own spare name, or refused by name - beforeoverride_invented_playersruns over the same field (which the golden then shows finding nothing to do). Each replacement is a recorded decision. A router name the question spells differently - a bare "Jokic" left bare, "Jaylen Tatum" for the question's "tatum", "Deron Williams" two edits from "derozan" - takes the question's own resolved name (Subject.named), which was_question_derived_player's job. - Decisions as data:
query/decisions.py'sDecision(stage, field, before, after, reason)records what was decided on the way to an answer - a reading of the question, an override of a routed field, a default - kept on the history record (adecisions:section of JSON lines beside the trace), onAnswer.decisions, on the API'sAnswerResponse.decisions, and shown by the page under a "decisions" disclosure; never parsed back out of the trace. The subject reading is the first producer: every fast-path answer now records who the question was read to be about, beside the repair chain that still writes the slots. query/subject.py: one reading of who a question is about (Subject,read_subject,question_supports), made from the question's own spans with the router's slots as a hint only where the question supports them. Measured before it was written: agrees with the fast path's ten-step repair chain on 279 of 290 recorded questions, and is right on the four where the chain is wrong. Nothing reads it yet - the roadmap's plan item 1 wires it in step by step against the golden.team_outlook's card on the page no longer repeats its own boxes as notes (only the BPI's offense/defense split, the strength-of-schedule rank and the other snapshots remain beneath it), and a postseason snapshot's second record is labeled "regular season" (regular_season_wins/losses) rather than "projected" - ESPN's projection columns hold the finished regular season once the playoffs are on.- "Vs southeast division", "against eastern conference teams", "in the west": the router keeps the conference or division's name with the word as the
situation, so the relation reads it (it captured the bare word "division", which the alignment reader refused by name - #213, yardstick-v2 F055). A bare word still captures and is still refused honestly. The question then reached the compiler as aplayer_splitsline over a boolean measure, which builtAVG(<boolean>)and crashed in DuckDB: a boolean measure on a one-figure intent is now its count (8 road double-doubles against the division, the key), andcore._aggrefuses to average or sum a condition at all rather than send it to the database. - Saving an edited note again replaces the earlier one in the history file instead of appending a second:
POST /api/notesreturns the line it wrote, and a save passing it back asreplacesrewrites that line in place (append_note(replacing=...); a line no longer in the file appends, so a correction is never lost). The page's note box holds the line its last save wrote and passes it back, so editing a note and saving again is a correction, not a second note. - "Who attempted the most three pointers" ranks attempts, not makes (the question's own word decides the column), and a championship question is refused before a team ranking answers it fluently and wrongly - titles are not on record as such.
- The web page shows every rendered answer's sentence first and at reading
size (it was the muted 12.5px caption), shows what the sentence goes on to
say beneath the table - a footnote's key, a coverage caveat, how a name was
read - instead of folding it under the text toggle (data["notes"], filled
by the templates andagent.pyas they attach them; the compiler's notes
are lines of their own), draws the sparkline at full width, folds a
season-long game table under its sparkline, renders a table-less answer as
prose rather than monospace, formats a rate column as a percent, and draws a
composed answer by its skeleton (rows, grouped, scalar) whatever intent the
router filed.team_record's month tables,period_leaderboardand a
player log's venue, result and per-game line render too. scripts/preview_answers.pyrenders any recorded question through the page
without the router - from a live-run jsonl, a history directory or an
explicit case - and screenshots each answer in headless Chromium.- The web page's tables are typeset in the page's own proportional font
(font-variant-numeric: tabular-numskeeps digits aligned) instead of
monospace, with a header rule and row separators; a percent column now
right-aligns like any other number. A chart-only answer (fingerprint,
shot_chart) shows a headline and its notes for the first time - it used to
render nothing above the iframe but the artifact itself, with which players,
which season and a substitution note ("only one of them matches anybody in
the warehouse") visible only under the collapsed text toggle. A player's
game log renders the extra columns a question asked for (data.columns,
e.g. FGM/FGA) instead of always the same four; a single-leader question
("who leads the league in total netpoints") still renders as a ranked table
instead of falling back to monospace text; a rate column whose label spells
out "percentage" or "pct" (a leaderboard's own wording) formats as one; a
margin/differential column keeps its sign;player_stat's per-game figure
and season total no longer share one label ("PTS" twice);period_split
folds a long per-game table under its sparkline the wayteam_quarter_points
already did, and a sparkline needs three points, not two, to draw. A
redundant caption (equal to, or a prefix of, the headline) no longer shows
twice, and a note re-derived from an answer's own sentence no longer repeats
underneath a body already built fromdata-data.notesalone once a
renderer has drawn something, sinceanswerNotes's text-parsing fallback
could only guess at the same sentence a structured body already read
correctly. A numeric table column now shows one decimal throughout once any
cell in it has one (table()'spad), except a synthesized summary row (a
game log's "per game" line,skipPaddingRows) or a transposed table where a
column mixes a count with an average (player_compare,player_matchup,
pad: false), which format their own decimals per row instead. - A ranked table (leaderboard, threshold_count, period_leaderboard) shows its
top 10 rows and folds the rest under their own count instead of listing
every one - a "top 50" question drew all 50 rows regardless of what was
asked. player_historyfolds a season table over 12 seasons ("past 20 years") the
same way a long game log does, and a sparkline is only drawn from points
that are real finite numbers - a mismatched key between the data and what
the chart reads it under used to draw a line throughNaNcoordinates
(see ISSUES.md, "Template data the page cannot render from",
player_history/twoPointFieldGoalPct) rather than simply not drawing one.- The web interface fits a phone. Below 600px the coverage pills are one row that scrolls sideways, and the header, footer and answer padding tighten - the chrome took about 80% of a 667px-tall screen and now takes under a quarter. A column heading wraps rather than pushing the next column off-screen (a comparison's second player started past the edge), and a table wider than its answer shows a shadow on the side with more to scroll. On a touch screen the question box is 16px, so iOS no longer zooms into it, and the keyboard hint is hidden; the page is sized to
100dvhand clears the notch and home indicator. Shot charts and fingerprints scale down to the frame they are drawn in (svg { max-width: 100% }) instead of being clipped, with the fingerprint's skill labels enlarged so they stay legible; at desktop width the page is pixel-identical. - The web interface shows whether it is connected - a dot in the top bar, green when the server answers, amber after one missed check, red after three - and reloads itself when the server restarts (a new commit or a rebuilt warehouse, where the server is deployed to restart on those), but only when that loses nothing: with an answer on screen, a draft typed or a question in flight it says "updated · reload" instead, and a click reloads. The page polls the new
GET /api/ping(association.web.app.PingResponse): an instance id drawn fresh each time the server starts, and whether it is answering. Unlike/api/healthit opens no warehouse connection and never probes ollama. Answer.history_filenames the.history/record an answer was written to (set byAgent.askin the same step that writes it); the web runner reads it off theAnswerrather than parsing the[history] ...trace line, which the web path carries verbatim.- The web interface can annotate an answer: a note control un...
association 4.3.0
- The docs describe the shared relation. The architecture page, the
templates package docstring, the README and the usage guide now say that the
box-score templates composequery/player_gamesrather than each writing
its own SQL, and their examples of narrowings no template honors no longer
name "under 14 FTA" and "since 2020", which both answer now. - "his last game" is one game. A
player_statquestion naming a single
game at one end of the span ("his last game", "her first game of 2026")
getsorderandlimit: 1set together inroute()from the question's
own words, andplayer_stathands it togame_log; "show maxey's stats for
his last postseason game" answers that game rather than his postseason
average. A fillerorderon a question naming no such game is still
dropped, as before (closes #142). - An unscoped count by a named player reads as his career. Product
decision (2026-09-19): "how many times has embiid fouled out?" is 0 this
season and 9 in his career, and only the second is the question.route()
sends athreshold_count"how many" question with a player and no season
in sight to his career; a season the question names ("this season", a
year, "his 18th season") still wins, and a league-wide count keeps the
default season. The answer names the scope it used, as it always did. - "Stats vs X" ends with the meetings behind the average. Product
decision (2026-09-19): averages over every meeting in scope, the game
count, and a short footer of the meetings themselves, newest first
(RECENT_MEETINGS, 5) - date, venue, result, points, rebounds, assists.
"Jayson tatum stats per game vs sas" now says "in 1 game" and shows the
only meeting beneath it. A per-game log is stillgame_log's, for a
question that says log, each game or last N. - "This postseason" names the current season, like "this season" does:
"maxey's stats for game 4 against the knicks this postseason" used to read
as a career question and ask which Maxey. And a stat's own name is never a
count's subject: "the most 30+ point 10+ rebound games" read "point" as a
player and answered for Sir'Dominic Pointer. - "His 18th season" is a season, settled once the player is known. "how
many 40+ points games does lebron james have in his 18th season?" arrived
as season 2018 - the ordinal read as a year - with LeBron dropped, and was
answered as the 2018 league leaderboard; then it refused as asituation.
route()keeps the ordinal asseason_nand drops a year the question
itself never named;game_log,player_statandthreshold_countresolve
the player over his career first and then settle the ordinal against his
regular seasons on record (templates.common.settle_ordinal_season),
answering "in his 18th season (2021 regular season)". A player with fewer
seasons than the ordinal is told how many he has rather than answered for
his last one, and a league-wide count ("most points in 15th season played")
refuses: the ordinal is a place in one player's career.
The same question's subject, dropped by the model, is restored from a third
grammar ("does lebron james have", beside #148's two), so it answers his one
40-point game of 2020-21 rather than falling through. - "Game 4" is one game of each playoff series. "Ayton stats in game 4
playoff games" answered with his whole postseason, then refused as a
situation, and "show maxey's stats for game 4 against the knicks" sent
the agent into a six-minute loop (#145).route()reads "game N" (1-7,
"game 7s" included - it is no longer around) into agame_nslot;
the relation numbers every postseason game within its series by date over
real_games(the series' own games, so a game he sat out still counts
toward the number) and keeps the Nth.game_logandplayer_stathonor
it, saying "in game 4 of the series" with an opponent and "of each series"
without; a regular-season question refuses, since nothing there is game 4,
and a team's log refuses it for now. - A line on a box-score stat keeps a player's games. "Sga games with
under 14 fta in his whole career" used to be refused (belowwas a slot no
template honored) and, before that, answered as 14 or MORE free throws
made - the model's nearest stat, the other way round.route()now keeps
the words after the number ("under 14 fta", as a list, one per phrase) and a
newaboveslot carries a minutes floor ("with 25 minutes", "20+ mins",
formerly a refusedsituation);templates.common.measure_filtersreads
them onto theplayer_gamerelation (MEASURE_WORDS: "fta", "fga",
"mins", "threes", ... to a column) and refuses a word it cannot map rather
than filtering on a guess.game_log,player_statandthreshold_count
honor both, and the answer says what it kept ("with under 14 free throw
attempts"). Athreshold_countphrase carrying the count's own number is
that count misread, so the phrase wins; one with another number is a second
line beside it. A team's log refuses a line, since it has no such column. ISSUES.mdkeeps one heading per priority tier, and a gate says so.
scripts/check_issues_md.sh(pre-commit and CI) requires exactly one
## P1:..## P4:heading, in order. An edit deleting the last entry of
the P1 section had swallowed the## P2heading with it (898ef66), so every
P2 read as a P1 andsync_issues.pywould have labeled a new one that way;
#127 had the P3 heading missing for days before that. Restored here.- A filler
limitonplayer_statis dropped whatever its size. Now
that a limitedplayer_statis answered as a log, a count the question never
named would answer a different question: "Portis vs bulls 2019-20 to
2023-24" arrived withlimit: 5and became a three-game log where his
averages were asked for.route()drops any limit onplayer_statthat no
count word supports, not only a 1; a year is not a count, and neither half
of "2019-20" is. "last 5 games" and "top 10" keep theirs. - A nickname another slot holds is not the subject. "myles turner bucks
stats without giannis last 10" routed withwithout: ['giannis'], and the
nickname override, seeing exactly one nickname in the question, rewrote the
subject to Giannis - who then could not play without himself.
override_nicknamesnow leaves alone a nickname the router already put in
another slot; the question answers with Turner's last 10 games without him. sinceis a scope the relation honors. "jokic vs cade since 2022"
answers their 7 meetings across 2022-2026 instead of this season's one:
_span_ofand_condition_scopetakesinceand build a span from that
season on (the phantom still excluded, never earlier than the table
reaches),game_log,player_statandplayer_matchupdeclare it in
HONORED_SCOPING, and a since-span's answer says "since 2022" rather than
"over his career". Defined once on the relation's span builders, so any
template on the relation gets it by declaring it.- A player's numbers over the last N games are the log with its averages.
player_statwith alimitor anorderhands the question togame_log
instead of refusing it - the product decision that "stats over his last N
games" is a per-game log with averages beneath, never the season line.
player_statdeclaresorder. Note the dependency this creates: a filler
limitthe router emits on a question that named no count is dropped in
route()(_LIMIT_REFUSING_INTENTS); the template now trusts the slot it
is given, so that rule is what stands between "westbrook stats as a
starter" and a one-game log. - The unseen-games counters have one home.
player_games.scope_without_guard
is the span clause for the reads that count what the played guard drops
(_empty_box_scores,_rebuilt_in_scope), replacingtemplates.players._box_scope;
the relation's module docstring says why those reads omit the guard and
where the team-level counterpart lives. With this, "a player's games" is
defined once (closes #149). - The last player-games read is on the relation.
conditions._player_games
(behindstreak,player_splits,record_whenandwith_without) now
derives its join, its played guard and its rebuilt-line blanking from
query/player_games.pyinstead of its own join toreal_gamesand
team_box_stats- the side and the scores come fromgamesdirectly.
Measured first: on the 2026-09-19 warehouse no played box-score row sits on
a gamereal_gamesdrops (43,504 games, 43,353 real), so the two reads
agreed by fact; now they agree by construction.player_splits' venue and
opponent narrowing names the relation's columns for a player's games and
keeps the team-box spelling for a team's. Pure refactor: all 11 recorded
cases across those four templates and all 195 across the five player-games
templates are byte-identical before and after. - The pair relation.
player_games.paired_rows_sqlreads the games two
players both played, on opposite teams or the same, as two reads of the
player-games relation joined on the event - both under the played guard,
with the rebuilt-line blanking applied byplayer_games.column.player_matchup's
meetings now come from it instead ofconditions._meetings' own join to
real_gamesandteam_box_stats. Pure refactor: all 17 recorded matchup
cases (195 cases across the five player-games templates) are byte-identical
before and after. - One definition of "a player's games."
query/player_games.pyis the
relation every box-score template reads through: the season-keyed join to
games, the phantom-1993 exclusion, the did-not-play and empty-line guard,
the rebuilt-line opt-in and the teammate tenure rule live there once, with
three skeleton readers (rows, one aggregate, aggregates per group)...
association 4.2.0
- A NetPoints name ESPN spells with a generational suffix, or hyphenates
differently, now matches too.match_keyreduces both sides to a
comparable form in four measured steps - diacritics dropped, hyphens to
spaces, whitespace collapsed, and a trailing suffix from a closed set
removed - which recovers 710 of the 1,060 remaining unmatched per-game rows:
Jimmy Butlerto ESPN'sJimmy Butler III(498 rows, his whole per-game
record),Rondae Hollis-JeffersontoRondae Hollis Jefferson(144),
Trey Jemison,Billy Garrett,Darius Brown. The suffix step is the one
that can merge two real people, and it does - ESPN holdsGary Paytonand
Gary Payton II,Tim HardawayandTim Hardaway Jr., fathers and sons, 31
colliding keys in all - so a reduced spelling more than one athlete id can
reach is never usable, and those names resolve only exactly. What is left
unmatched is a different name rather than a different spelling (a nickname, a
short first name, a middle name, a reversed order) and wants a curated list. - A NetPoints name spelled with diacritics now matches ESPN's spelling of
it. NetPoints' 2026 files sayNikola Jokićwhere ESPN'splayerssays
Nikola Jokic, so an exact match lost those players their whole 2026 season:
Jokic matched every season from 2019 to 2025 and none of 2026, and
re-fetching cost 682 already-resolved player-games. The name coming in from
NetPoints is now folded (NFKD, combining marks dropped) when the exact
spelling misses, which resolves 20 of the 39 unmatched names. Checked rather
than assumed: no ESPN name carries a diacritic (0 of 3,080) and no two fold
to the same string, so the fold cannot reach a player it was not already
about, and the exact spelling always wins where both exist. Only diacritics -
not case, punctuation or whitespace. scripts/backfill_netpoints_names.pyre-parses the NetPoints tables that
are matched by display name, without refetching ESPN. The 4.1.0 parser fixes
(#22, #101) mean the Parquet on disk was written by the old code, so a load
alone changes nothing - but a fulldata pull --forceover the NetPoints era
also refetches about 11,000 ESPN game summaries the fix does not touch, some
40 minutes at the default rate limit. This runs the NetPoints steps of a pull
and nothing else, through the samePipelinemethods,_write_rowsand
warehouse.build, and reports both measurements before and after: unmatched
rows per table, and the row counts of the players ESPN files under two ids,
which is the only one that shows #101 moving (that table drops an unmatched
row rather than nulling its id, so its unmatched count is always 0).
association 4.1.0
- The new pin rewriting refused its own first real bump, and now does not. Its
post-rewrite check greps the whole tree for the old pin, so it fired on a
comment insidebump_version.py- and would have fired onCHANGES.mdat
the next release, because a released entry quoting the tag it shipped under
is a stale pin on purpose. The check honors the same exclusions the
discovery does, and a test covers the case. - The version directives on this release's new public symbols name 4.1.0, not
4.0.2. Five agents working in parallel were each told 4.0.2, correct for a
release of fixes;GET /api/coverageis a new endpoint, which makes this a
minor bump. Corrected in the pre-release audit, which is what that audit is
for - and the pinned install commands were rewritten by the bump script
itself this time (#60), rather than by hand minutes before the tag. - The web page now reports
GET /api/coverage, what the warehouse actually
holds grouped into the three tiers a question can land in - box score, box
score plus play-by-play, and both of those plus NetPoints - and the page
renders it as a row of pills under the header. This replaces/api/health's
single min/max season as the page's only claim about coverage, which read as
"the whole span is answerable" when only its narrowest table was (#71): a
2016 shot chart and a 2016 NetPoints fingerprint looked equally reasonable to
ask for, and only one of them was. Each tier's floor is read straight from
association.nba.coverage.COVERAGE- the same enforced table a template's
own refusal reads - rather than recomputed from row counts, so the page and
a refusal cannot say two different things about the same season; only how
far a tier's data currently reaches (last_season) is counted live, since no
floor records that. A season inside a tier's range but declared partial
there (2002's play-by-play, 2002-2003's shot chart) is marked rather than
shown as uniformly whole, and 1993 - ESPN's phantom copy of 1994 - is
excluded from the box score tier's range rather than offered. - A shooting-percentage leaderboard's qualifier now scales to a shortened
season instead of applying an 82-game-calibrated floor flat.ts_pct
(550 true-shooting attempts),efg_pct(480 field-goal attempts) and
fg_pct(400 field-goal attempts) were measured against the warehouse: at
the flat floors, 2020 qualified 157 players and 2021 155, against 174-184
in every full season measured (2019, 2022-2026), and the 66-game 2012
lockout season qualified 127 - a published rule scales per team game, so
this was a stricter qualifier than the one it claimed to be, not a missing
one.LeaderboardMetric.scales_with_schedulemarks the three floors this
applies to;leaderboard.default_min_samplescales them from the season's
own MEDIANreal_gamesteam-game count (72 for 2020 and 2021, 66 for
2012), rounding half up, and falls back to the flat floor when
real_gamesis unavailable. Re-measured: 2020 moves to 171-183 across the
three metrics, 2021 to 176-220, 2012 to 176-211, and every full season is
unchanged (proven identical, not just close, since the scaling factor at
82/82 is 1).LeaderboardResult.min_sample_applied- already wired into
the answer text - now names the real, scaled number rather than the flat
one.three_pt_pctandft_pctshare the identical shape and are not
scaled yet - see ISSUES.md. - NetPoints per-date rows that fail the exact display-name match now keep
the source name instead of dropping it.parse_net_points_dailyand
parse_net_points_daily_playersused to writeathlete_id=Nonewith
nothing else on the row, so a query grouping by(event_id, athlete_id)
counted every unmatched row on a date as a duplicate of every other -
2,190 rows innet_points_player_gameand 64,210 in
net_points_player_game_fingerprint, measured 2026-09-18. Both tables now
carrydisplay_nameon every row (matched or not), and
net_points_player_gamealso carries NBA.com's ownnba_player_id
(plyrIDin the source), where the file provides one (ISSUES.md #22). - A NetPoints display name shared by exactly two locally-known athlete ids
is no longer dropped unconditionally.Pipeline._name_to_athlete_idused
to treat any name with more than one match as permanently ambiguous, which
meant every one of the 8 players ESPN files under twoathlete_ids in the
same box score (#87) had zero rows innet_points_player_game,
net_points_player_game_fingerprintandnet_points_player_fingerprint
for their entire career - Corey Brewer, 985player_box_statsrows over
2008-2020, had no per-game NetPoints for a single one of them. A new
Pipeline._resolve_duplicate_athlete_pairsapplies the same proof
fetch/repairs/duplicate_athletes.pyuses to merge those ids at load time -
the pair appears in the SAME team's box score for the SAME game - straight
off theplayer_box_statsParquet tree at fetch time, and resolves the name
to the established id when that proof holds. Measured against the live
warehouse: all 8 known pairs now resolve, to the identical id
player_box_stats_dedupedalready treats as canonical, and none of the
other 13 names shared by exactly two ids (real different people, #21) picked
up a false match (ISSUES.md #101). scripts/bump_version.pyrewrites the pinned install commands itself, so
a release can no longer ship instructions that install an older one. The pins
saidv1.4.0through three later releases, 3.0.0 shipped still pointing at
v2.2.0, and 4.0.0 and 4.0.1 were only right because a human edited them
minutes before each bump. The script discovers them withgit greprather
than a hardcoded list (a list of two files is what letdocs/usage.rstrot
unnoticed), refuses to run if a pin names neither the current nor the new
version, and asserts no old pin survives the rewrite. It skipsCHANGES.md,
tests/andscripts/, which hold the same pattern as release history,
fixtures and the regex itself - scanning them refused a real bump, which a
test against this repository now catches.scripts/bump_version.pynow rewrites the install-command pins itself
(#60). Because PyPI is unreachable,README.md,docs/installation.rst
anddocs/usage.rstpin a release tag
(git+https://github.com/jeffknupp/association@vX.Y.Z), and nothing rewrote
them automatically - they saidv1.4.0through three later releases, and
3.0.0 shipped still pointing atv2.2.0. The bump script now finds every
such pin withgit grep(not a fixed file list, so a pin added to a new doc
is covered the same way), rewrites@v<current>to@v<new>in the same
run as the version bump, refuses outright if a pin names a version that is
neither the current one nor the new one, and asserts no@v<current>pin
survives anywhere in the tree afterward.docs/releasing.rstnow describes
this instead of telling a human to update the pins by hand.- A retired player's question that names no season no longer defaults to
the current one and stops there.player_stat,single_game_high,
game_log,player_netpointsandshot_chartall read a missingseason
slot as "now", and a retired player's "now" is empty - "Allen Iverson's
points" answered "Allen Iverson has no 2026 regular season numbers in the
warehouse", true and about a year nobody asked for. The refusal now
redirects to what the warehouse actually holds for him when the season was
defaulted rather than named: "... He last appears in 2010. The warehouse
holds his 1997-2010 regular seasons; name one, or ask for his career." Never
substitutes an answer, only names where to ask again - the same discipline
entities.suggest_playersalready follows for a near-miss name. A season
the question names outright keeps its plain refusal, because that answer is
correct as given:_Spannow carries adefaultedflag from_span_ofso
every reader downstream can tell the two cases apart.player_netpointsand
shot_chartdo the same from their own local flag, since neither is built
on_Span; neither offers "or ask for his career", since neither template
has a career span to redirect to. (#18)
association 4.0.1
- A narrowed
game_logorplayer_statquestion over a season whose box
scores ESPN served empty no longer says the games do not exist. Both read
_no_narrowed_gameswhen their guard leaves nothing, and it used to check
only for a recorded (or rebuilt) box score - so a stat outside
REBUILT_STATS(turnovers, fouls, 3PM,plusMinus) sent the read back to
the fetched lines, which are empty for every 2013-2018 Chicago or New
Orleans game, and the answer became "No 2015 regular season games found for
Anthony Davis" of a man who played 68. It now checks whether the games exist
with an empty box score before saying they do not exist at all - "Anthony
Davis played 68 games in the 2015 regular season, but the box score is empty
for all of them" - the mirror-image bugAGENTS.mddescribes, in
single_game_high's own shape. player_statnarrowed byopponent,venueorwithoutnow reads a line
rebuilt from play-by-play in place of an empty ESPN box score, for a stat the
rebuild gets right. It never did before, even for points - the one stat
measured most accurate - so "Anthony Davis points vs the Lakers in 2015"
answered the same wrong-cause refusal as a stat the rebuild does not trust.
game_logalready read rebuilt lines for its own narrowed span; this brings
player_statin line with it. A shooting percentage still never widens to a
rebuilt line - its attempts are outside what the rebuild was measured for.- A power-index answer whose snapshot carries no rating now says so, rather
than dropping the line. The power index is that answer's headline, so the old
behavior left a reader with a record, a projection and chances and no sign
that the number they asked for was missing. ESPN's 2026 regular-season
snapshot is the live case - all 30 teams NULL inbpi,bpioffenseand
bpidefensewhile their records and projections are populated, and the only
one of the table's 21 season/season-type groups with any NULL rating - and it
is the snapshot a 2026 regular-season question now reads. - A regular-season BPI question now reads the regular-season power-index
snapshot outright, instead of whichever pre-playoff snapshot ESPN stamped
last. Once the paging fix gave every snapshot all 30 teams, the play-in
snapshot (season type 5) started postdating the regular-season one in 2023,
2025 and 2026, so the same question named a different snapshot depending on
the season - "how good were the Knicks in the 2026 regular season" answered
from the play-in view.team_outlook(query/templates/teams.py) now finds
theseason_type == 2snapshot directly and only falls back to the latest
other pre-playoff snapshot, then the postseason one, where no
regular-season snapshot holds the team - true of no season the warehouse
holds today (2017-2026), measured read-only. SeeISSUES.md, "A
regular-season BPI question answers from the play-in snapshot in 2023, 2025
and 2026" (#88). ESPNClient.get_collectionnow warns when the first page of a collection
cannot be read at all, instead of quietly returning[]._request_json
returnsNonefor a 400 or 404, which used to hit
if not isinstance(data, dict): breakwithexpectedstillNone, so the
existing declared-vs-fetched warning had no count to compare against and
never fired - an endpoint that started rejectinglimit=1000with a 400
would have reproduced the original 25-row power-index bug with nothing in
the log. A 200 whose body is not the paged-collection shape (noitems
list) is the same failure and now warns the same way. A later page failing
is unaffected: page one'scountis already on record by then, so the
declared-vs-fetched warning already covers it, and a genuinely empty
collection (count: 0, one page, no items) still logs nothing.- The web page's games count no longer counts rows that are not games.
gamescarries 151 of its 43,504 rows that were never played - placeholders,
team-slots naming an id no franchise has, phantoms and a duplicate - so the
health line said 43,504 where 43,353 were played. It readsreal_gamesnow,
falling back togamesfor a warehouse loaded before that view existed,
which is the same checkconditions.box_sourcemakes for the filled box.
association 4.0.0
- The version directives on this release's new public symbols name 4.0.0, and
the pinned install commands inREADME.md,docs/installation.rstand
docs/usage.rstpoint atv4.0.0. Both are pre-release corrections rather
than changes: six directives said 3.1.0, written before the breaking change
below settled the number, and 3.0.0 shipped with its pins still on v2.2.0. - A coach question is refused, naming the real cause, instead of falling
through to the agent. No table here holds a coach, so the agent queried
tables with no such column and was then free to fill the silence from its own
weights - the failurecheck_coverageexists to stop. The refusal says what
is actually wrong rather than blaming the source, because "ESPN does not
publish coaches" was probed and is false: it serves two coach collections and
neither is usable (the season-by-season one ignores the season it is asked
for and returns today's staff, the per-team one covers 12 of 30 teams in
1996, never shows a mid-season change, and names the wrong coach for some
franchises outright).route()assigns the intent from the question's own
words, soROUTER_PROMPTandROUTER_SCHEMAare untouched - both hashes
unchanged - and no other question's routing can have moved. - A player ESPN lists twice in one team's box score, under two
athlete_ids, is now merged into one row. Measured against the
2026-09-17 warehouse: 8 players, 69 team-games, none before 2003. A new
player_box_stats_dedupedtable (association.fetch.repairs.duplicate_athletes)
picks the id with more career games carrying real minutes and keeps the
real row over a fabricated all-zero blank; a pair is merged only where
every shared game is safe (one side has no minutes, or both sides agree
exactly), so a future pair that genuinely disagrees is left unmerged rather
than guessed at.player_game_lognow reads the merged table, so a
per-game lookup for one of these 8 players (single-game highs, streaks) no
longer sees a fractured career under two identities. Team-level sums and
player_advanced_statsstill read the raw table and are unaffected - see
ISSUES.md. - A team-level question about Vancouver's 1995-96 season is no longer
answered from NULLs it did not have to be. ESPN serves an all-NULL
team_box_statsrow for every 1995-96 Grizzlies game - and, unrecorded
until now, for 25 more teams' games against them that season, plus one 2000
game - while itsplayer_box_statsrows for the same games are real, with
real minutes.team_box_repairnow rebuilds field goals, three-pointers,
free throws (with their shooting percentages), assists, steals, blocks,
fouls, individual turnovers and the offensive/defensive rebound split from
those player rows on any team row this shape touches - measured exact on
all 2,387 surviving 1996 regular-season team rows and all 2,472 surviving
2000 ones.totalReboundsand the columns the player box has no sibling
for (team turnovers, technicals, flagrant fouls, points in the paint, a
largest lead) still have nothing to rebuild from and stay NULL. - A team's rebounds are now comparable across the 2021/2022 season
boundary. ESPN's team boxtotalReboundsstopped counting rebounds it
credits to no player from 2022 on, so a split spanning the change (or any
comparison of an old season with a recent one) showed a team's rebounding
falling off a cliff for no basketball reason - measured at ~52 a game in
2019-2020, ~48 in 2021, 44.45 from 2022 on.player_splitsandstreak's
team-rebounds reads (query/conditions.py's_team_gamesand_TEAM_LINE)
now useoffensiveRebounds + defensiveRebounds, which already equals ESPN's
owntotalReboundsin every season from 2022 on and is populated everywhere
totalReboundsis.team_metrics.py'savgRebounds(used byteam_stat
andteam_leaderboard) was checked and needs no change: unlike the
game-level column, ESPN's season aggregate already equals
avgOffensiveRebounds + avgDefensiveReboundsin every season since 1994. - A playoff answer from 1995-1998 now says which games have no box score.
ESPN lists eight of those postseason games and serves each one a box score
with no player lines in it - probed live, all eight return aboxscore
carrying zero athlete lines where control games in the same seasons return
24, the athlete gamelog omits them, andplaysstarts too late to rebuild
them. 1997 is the one that costs an answer: the whole Chicago-Miami
conference final, so Michael Jordan's postseason read 14 games and 439 points
against ESPN's own 19 and 590, with nothing said. The game LIST is
deliberately not caveated - those games are in it, with scores and a winner,
so a playoff game count or head-to-head record over them is right. - Breaking:
Coverage.partialandpostseason_partialnow map each season
to its own sentence, replacing a tuple of seasons plus one shared note;
partial_noteandpostseason_partial_noteare gone.team_box_statsis
short five games of the 1997 playoffs and ten of 2001's for unrelated
reasons, and one shared note named both in an answer about either - the
wrong-cause noise that module exists to stop. Nothing outside
association.nba.coverageread the two removed fields, andseason in coverage.partialstill works, so the change is breaking only for a caller
that read a note off aCoveragedirectly. The next release is therefore
4.0.0, which is what the version directives here name. - A 2013-2018 shooting leaderboard now says who is missing from it. Those
rates are summed from the box scores ESPN serves zeroed for every Chicago and
New Orleans game, so 21 to 33 players a season clear the qualifying floor by
ESPN's own season totals and fall under it in the advanced table - 153
player-seasons, of which 83 are on the other 28 teams, short only the games
they played against those two. The 2015 true-shooting board omits Tyson
Chandler, its runner-up. The seasons are declared partial rather than filled
from the rebuilt box: a rebuilt season total is exact about half the time and
biased low, and no attempt column was ever measured. - A 2001 playoff question answered from the player tables now says what is
missing. Ten games of that postseason are not in ESPN's archive anywhere,
and the caveat saying so was declared only ongamesandteam_box_stats-
whichsingle_game_highandthreshold_countnever read, so "Shaquille
O'Neal had 4 games with 30+ points" was stated as fact over 11 of the 16
playoff games he played.player_box_stats,player_game_logand
player_season_advanced_statsnow declare it too, in the player's own words:
40 of the 190 players with a 2001 postseason line are short in the box
scores, 152 games in all. ESPN's per-player season line is complete and is
deliberately left uncaveated, since it is what proves the box scores short.
association 3.0.0
- No function is more complex than radon grade C (cyclomatic complexity
20), enforced by a xenon gate in pre-commit and CI, with no module worse
than C and the average no worse than B. The worst were the router's
route()at 127,parse_game_summaryat 61 and a set of query templates
between 21 and 89; each is split into named steps called in the original
order. No behavior changes: every split was checked against the original
code by calling it with thousands of inputs (405,588 forroute(), 4,433
slot sets for the templates, every fixture for the parser) and comparing
the full results, besides the test suite. - Breaking: the package is reorganized, so the next release is 3.0.0.
Nothing about what the commands or the query engine do changes; module paths
do.association.cliis a package: the commands are in
association.cli.commands, andassociation.repo_pathsis now
association.cli.paths.association.cli:main(the console entry point)
andassociation.cli:cliresolve as before.- The modules both
fetchandqueryread moved from the package root
intoassociation.nba:association.season,association.coverage
andassociation.franchisesare nowassociation.nba.season,
association.nba.coverageandassociation.nba.franchises, and
association.net_points_categoriesisassociation.nba.netpoints. - The load-time repairs and filtered tables moved under
association.fetch.repairs:game_repair,team_box_repair,
season_totals_repair,reconstructed_boxandreal_games. association.query.templatesis a package, one module per subject:
common(the context and result types, the scoping and coverage checks,
and the helpers more than one subject uses),players,games,teams,
shots,netpointsandsplits. The package re-exportsTEMPLATESand
the scoping API, sofrom association.query.templates import TEMPLATES, check_scope, ...still works; a private helper is imported from the
module that defines it.ROUTER_PROMPT,ROUTER_SCHEMA,ROUTER_NUM_CTXand
ROUTER_PROMPT_TOKEN_BUDGETmoved fromassociation.query.routerto
association.query.router_prompt, byte-identical (hashed before and
after), so what the model is told sits apart from what is done with its
answer.
- Importing the web API no longer loads ollama.
association.web.runner
imported the queryAgentat module level for a type annotation, so
import association.web.appbrought in the model client that AGENTS.md says
the API layer must never import. The import is now type-checking only. The
rule and the package layering are enforced by import-linter in pre-commit and
CI:cli>web>query | check>fetch> the leaf modules,fetch
andqueryindependent, and nofastapi,uvicornorpydanticoutside
web. - Spelling is checked. codespell runs in pre-commit and CI with its
British-to-American dictionary, and the 37 findings are fixed - mostly
British forms in comments, docstrings and docs ("neighbouring", "cancelled",
"behaviour", "judgement"), plus "unparseable" and "pre-empts". Comments and
docs only: the router's compiled patterns and the prompt text hash
identically before and after. - Every imported package is declared.
botocore(imported by the
NetPoints client) is now a core dependency andpydantic(imported by the
web API) is in thewebextra; both used to arrive only throughboto3and
fastapi. Thedocsextra namespackaging, whichdocs/conf.pyimports.
Nothing new is installed. deptry now checks this in pre-commit and CI. - Dead code is now a gate. vulture runs in pre-commit and CI at its lowest
confidence, with the names only a framework calls listed, each with its
caller, invulture_whitelist.py. Its one real finding is removed:
TeamMetric.needs_opponent, set on four team metrics and never read since
the field was added. - Stricter lint, a dependency audit, and three pieces of dead code gone.
Ruff now also enforcesDTZ,BLE,RUF,PERF,C4,SIM,RET,PLW
andPLE; the 66 findings were fixed, and the handful that are deliberate (a
blindexceptat a boundary, the local calendar date incurrent_season)
carry an inline reason. No behavior changed:eastern_day_utc_rangenow
builds UTC-aware datetimes, which format to the same strings. pip-audit runs
overuv.lockin a new Dependency audit workflow on every push and weekly
(scripts/audit_dependencies.sh). Removed as unused:fetch.storage.write_row
(no caller outside its own test -Pipeline._write_rowis the real path),
fetch.team_box_repair.REBUILT_COLUMNS(never read; the repair's SQL names
its columns itself), and the privatetemplates._as_int. - The 2001 playoff caveat names every short series, and counts
Philadelphia right. A 2001 postseason answer said Philadelphia's run
"reads 15 games" and named only the LAL-PHI Final and the MIL-PHI
conference final. Finals Game 5, recovered by the scoreboard discovery pass,
brought Philadelphia to 16, and MIL-CHA (two games) and LAL-SA (one) are
short too. The note ongamesandteam_box_statsnow says ten games across
all four series and 16 against ESPN's 23. - The README's first install command works. It was
pip install 'association[web]', which the README's own note says does not
work while PyPI is unreachable; it now installs from the release tag, as
docs/usage.rstdoes too. The README also lists the 2001 playoffs' ten
missing games among its known limitations, and its project layout names
season.py,franchises.pyand the load-time repairs.
association 2.2.0
-
Docs brought up to date with the code. A pass over the README, the
Sphinx pages,AGENTS.md,DATA.md,ISSUES.mdand the package docstrings
for claims the last release's work had made untrue.docs/architecture.rst
now describes the load-time repairs (game_repair,team_box_repair,
season_totals_repair,reconstructed_box,real_games) and the
temp-file warehouse swap; the README anddocs/usage.rstlist quarter and
half questions, and the empty 2013-18 box scores are described as rebuilt
from play-by-play for per-game answers rather than only counted; the
development setup syncs thewebextra the tests need;docs/releasing.rst
says the PyPI upload currently fails. In the docstrings,
association.query.fingerprintno longer says a single game cannot be
fingerprinted,conditionsno longer describes a fixed five-hour Eastern
shift,team_metricsnamesreal_gamesas its source, and
override_invented_playersno longer says an invented name falls through to
the agent. Version directives that named 2.3.0 or 2.1.1 now name 2.2.0, the
release they ship in. -
The Pistons' 1990 title clincher is a Detroit win again. ESPN serves
Game 5 of the 1990 Finals (100614008, 14 June at Portland, Detroit 92-90)
with the two teams on each other's sides - Detroit at home, losing 90-92 -
so game logs printed the clincher as a home loss and the series read 3-2.
A new load-time repair,fetch/game_repair.py, puts the team ids back
(andteam_box_stats.home_away), keyed on the event id and guarded on the
stored row, so it is idempotent and stops by itself if ESPN corrects the
game. Checked across all 570 postseason series for impossible win counts and
out-of-format home games, it is the only such game outside 2001's missing
ones. -
Games stored with no tip time are dated the day they were played, not the
day before. ESPN stores such a game as midnight US Eastern -04:00Zin
summer - and every date this project printed or filtered on went through a
fixed five-hour shift, which is right for a real tip and moves a summer
midnight to the previous day. 391 games printed a day early in every game
log, single-game high and month split: 379 in 1988-1992, the whole
1989-1992 postseason among them (the Pistons' 1990 title clincher on 14 June
read 13 June), and 12 in the 2000-2001 postseason. A date question missed
them for the same reason. The fix is the real Eastern clock rather than an
era cutoff, because 2026 carries ten04:00Zstamps that are genuine 11pm
EST tips:association.seasongainseastern_utc_offset_hours,
eastern_date_sqlandeastern_day_utc_range, with the US daylight-time
rules written out so no tz database is needed, andeastern_datefollows
them. All six places that turned a stamp into a date -season, the
NetPoints matcher,real_games,conditions,team_metricsand the
templates' date filter - now use them, which also retires five duplicate
declarations of the offset (#82). Measured against the warehouse: exactly
those 391 games move, 318 of 318 whose event id encodes the date now match
it (0 did before), andreal_gamesis unchanged at 43,353 rows. (#76) -
2008's team rebound columns are rebuilt at load time. ESPN serves the
2008 regular season with the real offensive rebounds under
defensiveRebounds, the team rebounds underoffensiveRebounds, and a
totalReboundsthat counts the offensive boards twice - so "Celtics home and
away splits 2008" printed 60.8 rebounds a game. A refetch serves the same
values.fetch/team_box_repair.pynow rebuilds the splits from the player
box and the total as the players' rebounds plus the team figure, the
definition of the seasons either side; the 2008 postseason is clean and
untouched. (#74) -
A full warehouse rebuild now builds into a temporary file and swaps it
in, rather than replacing tables one statement at a time indb_path
itself.association data pull/load's full-rebuild path is exactly the
memory-hungry caseAGENTS.md("Working on the fetch path") describes -
playsalone from 17,500 files, on the same connection loading 17 other
tables - and building in place meant an interruption (an OOM kill or
anything else) left the tables already replaced at their new contents and
the rest at their old ones, with nothing recording that the build never
finished.fetch/warehouse._build_fullnow writes into
<db_path>.buildingand only replacesdb_pathonce every load, repair
and view succeeds; an interrupted build leaves the existing warehouse
completely untouched, and the leftover.buildingfile is itself the
marker the next full build logs and replaces. It also stops a full rebuild
from ever carrying forward a prior partial load's free space - the same 19
tables and views measured 1.73 GiB in a repeatedly partial-loaded file
against 0.92 GiB freshly built - since a full rebuild now always starts
from an empty file. A partial--tablesreload (used bydata pull's
incremental path, and by the backfill scripts) is unaffected: it still
writesdb_pathin place, because it depends on tables already there that
it is not reloading. -
A worktree can now run
association data pull/loadand the audit
scripts with no--data-dir/--db-pathat all. All nine call sites
(cli.py, andcheck_routing,check_coverage,check_nicknames,
check_net_points_games,check_team_box,backfill_season_totals,
backfill_missing_playoffsandbackfill_power_indexunderscripts/)
defaulted to the literal./nba.duckdband./data/parquet, which a
worktree does not have - both are gitignored build artifacts that live
beside the main checkout. New moduleassociation.repo_pathsresolves
each default to the current directory's copy where one exists, else the
main checkout's, found throughgit rev-parse --git-common-dir, else the
original literal default unchanged. Verified read-only from a worktree with
zero arguments:check_coverage.py,check_nicknames.pyand
check_team_box.pyeach ran and reported against the main checkout's
warehouse rather than failing with "database does not exist". -
A fresh worktree's venv is documented as needing a sync before the
gates run, and theCHANGES.mdgate now says so when it checks nothing.
AGENTS.md("Before you commit") gets the line CI runs -
uv sync --frozen --extra dev --extra docs --extra web- sinceuv run
alone creates a venv with none of thedev/docs/webextras and
uv run pytest -qfails before the suite starts. Separately,
scripts/check_changes_md.shread onlygit diff --cached, so it printed
"Passed" withsrc/edited but nothing staged - the same check that would
correctly fail once the edit was staged. It now says "nothing staged, so
nothing to check" and still exits 0, rather than reading as a real pass. -
Every team name an answer prints is the name it had that season. A 2005
Knicks log listed a game "vs Brooklyn Nets", eight years before the Nets moved;
the 2008 standings put the Charlotte Hornets 23rd, a team that did not exist
that year; every all-seasons streak answer ended "franchises are named as they
are today", which was an honest description of a bug. ESPN keys a team by
franchise andteamsholds only today's names, so every row joined to it
read today's.The franchise table moved to a neutral module,
association.franchises,
because both packages need it -fetch/warehouse.pybuilds the
player_game_logview's abbreviations, and the query templates name teams
everywhere else - and CLAUDE.md keeps those two from depending on each other.
It gained each era's abbreviation (NJ, SEA, VAN), and two renderings of one
lookup:season_namefor Python andseason_name_sql, a SQL expression that
names each row for its OWN season, which a career log crossing a relocation
needs. A test checks the two against each other at every season either side
of every boundary. Both rename only whenteamsfiles today's name under
that id.Applied at every place a team name is printed: team game logs and quarter
scores, home/road standings, team metrics, record tables, all-seasons streaks,
player-matchup meeting logs, with/without stints (a stint across a rename names
both, "New Jersey Nets / Brooklyn Nets"),period_splitrows, and the
player_game_logview. The SQL agent's example queries inprompt.pyare left
alone, because changing them spends preamble budget.The view needs a warehouse reload to show it - it is built at load time -
and was reloaded with this change. Every site was perturbed back to today's
name and watched to fail; four of the first five came back MISSED until each
had a test of its own, and the view's test had to move from 2005 to 1997 to
catch anything, because the Grizzlies were already in Memphis by 2005. -
The router prompt's size is documented correctly, and budgeted.
query/router.pysaid the prompt was "~430 tokens" in its published
docstring and besideROUTER_NUM_CTX; it is 9,989 characters, about 2,500
tokens at the four characters a token the agent's budget is measured at,
against a 4,096-token window. NewROUTER_PROMPT_TOKEN_BUDGET(three
quarters of the window) and a test that fails when the prompt plus a long
question passes it - the router's counterpart toPreambleTooLarge, as a
test rather than a runtime check because the prompt is a constant. Not
measured with the model's tokenizer; the figure is an estimate. -
A stale comment in
router.route()no longer credits the 2.0 REPL with
theprevious_questionfollow-ups; it now says what arrives there
(Agent.last_question, None in both shipped callers) and who the branch is
for. No behavior change. Thedocs/usage.rstassists example now carries the
"(mini...
association 2.1.0
The largest release so far on the query side. It adds eight new kinds of
question, and most existing ones can now be narrowed the way real questions
narrow them. The headline changes are below; each is detailed in its entry
further down.
- New question types.
- A team's numbers, team rankings, and a team's outlook (BPI, playoff and
title odds). - A player's splits: home and away, starter and bench, wins and losses, and
by month. - A team's record, or a player's line, with and without a teammate.
- A team's record when a player reaches a number.
- The games two players played against each other.
- Winning and losing streaks.
- A team's numbers, team rankings, and a team's outlook (BPI, playoff and
- Questions narrowed the way people ask them. A player's games and
averages against one team, at home or away, over a career, or without a
teammate. Career leaderboards for counting stats, averages and shooting
percentages (usage, true shooting, eFG% and NetPoints have no career
ranking). Career highs and career counts. - Names resolved from what the question says. A name the router invented,
dropped or completed is checked against the question. An ambiguous surname is
narrowed to the players who played in the season asked about, and every match
is considered, not just the first ten alphabetically. - Answers that were wrong, now right.
- True shooting and eFG% leaders qualify on attempts.
- Shot distances are measured from the rim, not from a point 5.25 feet
away from it. - A postseason before 1994 is found by the year it was played.
- ESPN's copied postseason lines are dropped.
- Every game is dated by the day it was played.
- Per-game NetPoints rows land on the right game.
- Narrow questions no longer get broader answers. Some filters no template
can apply yet: a playoff round, a season range, "under N", back-to-backs. A
question with one of these now goes to the slower agent instead of being
answered by a template about something else. A season the warehouse cannot
reach is refused, with the reason. - Four holes an outside review found are closed. Among them, the agent's
SQL connection can no longer read the disk, and the web server no longer
shares one conversation between every browser.
Known gaps and data faults are listed, ranked, in ISSUES.md. Among them:
every Chicago and New Orleans game from 2013 to 2018 but two has an empty box
score, and the 2000 and 2001 playoffs stop before the Finals.
-
A year before 1990 is read from the question. The question text and the
router'sseasonslot both discarded any year below 1990, a floor that
predatescoverage.py. So "who led the league in scoring in 1980" was
answered with the current season's leaders, and "Bulls record in 1985" with
their current record. The floor is now the league's first season (1947), in
one constant instead of two. A year below a table's first season reaches the
coverage check and is refused, with the reason. -
--include-net-points-dailysays what it fetches. Its help now names the
per-game play-type table it also writes, and says it makes two requests per
date, not one. -
An ambiguous name is narrowed to the season being asked about before
anybody is asked which one was meant. "How did curry do against the
celtics this year" answered'Curry' matches more than one player - did you mean Dell Curry, Eddy Curry, JamesOn Curry, Michael Curry or Seth Curry (1 others also match)?. The warehouse holds six Currys, the list was sorted by
name and the sentence names five, so the one it left out was Stephen - and
four of the five it did name never played in 2026. It now asksdid you mean Seth Curry or Stephen Curry?, the two who did.Every template that resolves a player (
player_stat,player_compare,
player_history,player_netpoints,game_log,shot_distance,
single_game_high,threshold_count,player_splits,with_without,
record_when,player_matchup,streaks) now narrows the candidates to
those with a row in the table its answer is read from, for the season or
span it will answer about - the rule charts already followed. A career has
every season in scope, so it eliminates nobody who ever played, but it names
whoever plays now first. It eliminates and never chooses: one survivor
is the answer because nobody else has a row to answer from, and two or more
are asked about, as Seth and Stephen still are. Three edges are deliberate.
A name matched in full is not narrowed, so "Gary Payton" in 2026 is still
told the father has no numbers rather than given his son's line. When
narrowing eliminates everybody, the question is asked exactly as before. And
player_historynarrows over every season up to the one it is anchored at,
since a history through 2026 still has Dell Curry's seasons to answer with.
player_netpointskeeps a candidate with a row in either of its two tables,
which disagree about who they hold (63 player-seasons are in the totals
only, 8 in the fingerprint only)._resolved_playernow requires the table,
so a template cannot resolve a name without saying where its answer comes
from - mypy refuses it.The cap was the other half.
find_playersreturns the first ten matches
alphabetically, and 71 name words match more than ten players ("Williams"
matches 62), so narrowing that page was choosing by alphabet - which the
chart path had been doing. In 2026's shot charts, 23 names drew one player
while others matching the name also had shots on a later page: "Davis" drew
Anthony Davis with JD Davison, Nigel Hayes-Davis and Trayce Jackson-Davis
eligible, and in the fingerprints "Brown" drew Bruce Brown with two more
Browns who had one. Narrowing now reads every match (find_players(..., limit=None)), on both paths and for the teammate a "without" names - which
was narrowed to his teammates over the same first page - and those names
ask. A clarification also
names every candidate from the season asked about rather than counting any
away (Ambiguous.active, passed toclarification) - 2026 has 14 players
surnamed Williams, and all of them are named - and a history names whoever
reached its last season first, so "Curry's scoring over the last 4 seasons"
lists Seth and Stephen ahead of Dell. What the "others also match" count
covers is now only ever players who could not be the answer, and it says "1
other" when there is one.Measured against the warehouse, over 6,325 player names, name words and
nicknames under nine template scopes: every one of the 5,349 that resolve to
a single player today resolves to the same player. Of the 976 that ask, for
2026 season lines 276 now resolve to the one candidate with a row, 191 ask
about fewer players, 484 have nobody in the season and ask as before, and 25
ask about as many or more - players the first page had hidden. No
clarification in any scope counts away a player from the season asked
about. Over thecheck_routing.pycorpus, routed once and answered before
and after through the agent's own slot pipeline, one answer changed - "What
was Curry's first game of the season?" still asks, now between Seth and
Stephen - and every name that resolved still does. Charts do change names
that resolve today, and on purpose: besides the 23 above, 8 names in 2026's
shot charts drew a best match with no shots while several matches had them
("Bob", "Marcus", "Scott") and now ask, and 4 draw the one match with data
instead of a best match without.The extra work runs only for a name that matches more than one player: 2-9ms
more for one season ("Curry" to "Williams") and up to 16ms for a history
span, warm, against a question that spends about 3s in the router. -
One player "compared" with a team is answered as his games against it.
"compare curry vs the celtics this season" arrived asplayer_comparewith
the Celtics as the second "player". Once they became theopponent,
player_compare- which reads season lines only - refused it, and the
question fell through to the agent whileplayer_statanswers it exactly. It
goes toplayer_statnow; two players and a team stay a comparison, and
still refuse the opponent rather than comparing whole seasons. Checked by hand
against the box scores: 22.5 / 5.0 / 6.5 over Curry's 2 games against Boston
in 2024-25, and for 2025-26 "none of them", which is right - he missed both,
inside a gap in his log from 31 January to 6 April.- A team in
playersis recognized by its LAST word being a nickname. The
player_matchupreroute matched one anywhere in the name, which sent
"magic johnson vs larry bird head to head" toplayer_stat- which refused
it - instead of the head-to-head template. All 30 team names end in a
nickname, and none of the warehouse's 3,080 player names does.
- A team in
-
The team a player's question plays against is the opponent, whichever
slot the router files it in. "compare curry and lebron vs the celtics" came
back withteam='Boston Celtics'beside the two players.scope_from_question
left it there - a team already inteamis howhead_to_headcarries its
own side - so noopponentwas set,check_scopehad nothing to refuse, and
nothing read the slot. With "steph curry" the answer was their whole 2025-26
lines, Curry's 43 games beside LeBron's 60, when Curry played none of them
against Boston. Where the template reads a player and one is present, that
team now moves toopponent, soplayer_statandgame_loganswer over
those games and every other template refuses.head_to_headand the team
templates are untouched. -
threshold_countanswers a surname only one player with games that season
has. "How many 30-point games did Curry have last season" asked which
Curry was meant even when only one of them played; the name is now narrowed
to players with a box score in the season a...