Repository navigation
association 2.2.0
-
Docs brought up to date with the code. A pass over the README, the
Sphinx pages,AGENTS.md,DATA.md,ISSUES.mdand the package docstrings
for claims the last release's work had made untrue.docs/architecture.rst
now describes the load-time repairs (game_repair,team_box_repair,
season_totals_repair,reconstructed_box,real_games) and the
temp-file warehouse swap; the README anddocs/usage.rstlist quarter and
half questions, and the empty 2013-18 box scores are described as rebuilt
from play-by-play for per-game answers rather than only counted; the
development setup syncs thewebextra the tests need;docs/releasing.rst
says the PyPI upload currently fails. In the docstrings,
association.query.fingerprintno longer says a single game cannot be
fingerprinted,conditionsno longer describes a fixed five-hour Eastern
shift,team_metricsnamesreal_gamesas its source, and
override_invented_playersno longer says an invented name falls through to
the agent. Version directives that named 2.3.0 or 2.1.1 now name 2.2.0, the
release they ship in. -
The Pistons' 1990 title clincher is a Detroit win again. ESPN serves
Game 5 of the 1990 Finals (100614008, 14 June at Portland, Detroit 92-90)
with the two teams on each other's sides - Detroit at home, losing 90-92 -
so game logs printed the clincher as a home loss and the series read 3-2.
A new load-time repair,fetch/game_repair.py, puts the team ids back
(andteam_box_stats.home_away), keyed on the event id and guarded on the
stored row, so it is idempotent and stops by itself if ESPN corrects the
game. Checked across all 570 postseason series for impossible win counts and
out-of-format home games, it is the only such game outside 2001's missing
ones. -
Games stored with no tip time are dated the day they were played, not the
day before. ESPN stores such a game as midnight US Eastern -04:00Zin
summer - and every date this project printed or filtered on went through a
fixed five-hour shift, which is right for a real tip and moves a summer
midnight to the previous day. 391 games printed a day early in every game
log, single-game high and month split: 379 in 1988-1992, the whole
1989-1992 postseason among them (the Pistons' 1990 title clincher on 14 June
read 13 June), and 12 in the 2000-2001 postseason. A date question missed
them for the same reason. The fix is the real Eastern clock rather than an
era cutoff, because 2026 carries ten04:00Zstamps that are genuine 11pm
EST tips:association.seasongainseastern_utc_offset_hours,
eastern_date_sqlandeastern_day_utc_range, with the US daylight-time
rules written out so no tz database is needed, andeastern_datefollows
them. All six places that turned a stamp into a date -season, the
NetPoints matcher,real_games,conditions,team_metricsand the
templates' date filter - now use them, which also retires five duplicate
declarations of the offset (#82). Measured against the warehouse: exactly
those 391 games move, 318 of 318 whose event id encodes the date now match
it (0 did before), andreal_gamesis unchanged at 43,353 rows. (#76) -
2008's team rebound columns are rebuilt at load time. ESPN serves the
2008 regular season with the real offensive rebounds under
defensiveRebounds, the team rebounds underoffensiveRebounds, and a
totalReboundsthat counts the offensive boards twice - so "Celtics home and
away splits 2008" printed 60.8 rebounds a game. A refetch serves the same
values.fetch/team_box_repair.pynow rebuilds the splits from the player
box and the total as the players' rebounds plus the team figure, the
definition of the seasons either side; the 2008 postseason is clean and
untouched. (#74) -
A full warehouse rebuild now builds into a temporary file and swaps it
in, rather than replacing tables one statement at a time indb_path
itself.association data pull/load's full-rebuild path is exactly the
memory-hungry caseAGENTS.md("Working on the fetch path") describes -
playsalone from 17,500 files, on the same connection loading 17 other
tables - and building in place meant an interruption (an OOM kill or
anything else) left the tables already replaced at their new contents and
the rest at their old ones, with nothing recording that the build never
finished.fetch/warehouse._build_fullnow writes into
<db_path>.buildingand only replacesdb_pathonce every load, repair
and view succeeds; an interrupted build leaves the existing warehouse
completely untouched, and the leftover.buildingfile is itself the
marker the next full build logs and replaces. It also stops a full rebuild
from ever carrying forward a prior partial load's free space - the same 19
tables and views measured 1.73 GiB in a repeatedly partial-loaded file
against 0.92 GiB freshly built - since a full rebuild now always starts
from an empty file. A partial--tablesreload (used bydata pull's
incremental path, and by the backfill scripts) is unaffected: it still
writesdb_pathin place, because it depends on tables already there that
it is not reloading. -
A worktree can now run
association data pull/loadand the audit
scripts with no--data-dir/--db-pathat all. All nine call sites
(cli.py, andcheck_routing,check_coverage,check_nicknames,
check_net_points_games,check_team_box,backfill_season_totals,
backfill_missing_playoffsandbackfill_power_indexunderscripts/)
defaulted to the literal./nba.duckdband./data/parquet, which a
worktree does not have - both are gitignored build artifacts that live
beside the main checkout. New moduleassociation.repo_pathsresolves
each default to the current directory's copy where one exists, else the
main checkout's, found throughgit rev-parse --git-common-dir, else the
original literal default unchanged. Verified read-only from a worktree with
zero arguments:check_coverage.py,check_nicknames.pyand
check_team_box.pyeach ran and reported against the main checkout's
warehouse rather than failing with "database does not exist". -
A fresh worktree's venv is documented as needing a sync before the
gates run, and theCHANGES.mdgate now says so when it checks nothing.
AGENTS.md("Before you commit") gets the line CI runs -
uv sync --frozen --extra dev --extra docs --extra web- sinceuv run
alone creates a venv with none of thedev/docs/webextras and
uv run pytest -qfails before the suite starts. Separately,
scripts/check_changes_md.shread onlygit diff --cached, so it printed
"Passed" withsrc/edited but nothing staged - the same check that would
correctly fail once the edit was staged. It now says "nothing staged, so
nothing to check" and still exits 0, rather than reading as a real pass. -
Every team name an answer prints is the name it had that season. A 2005
Knicks log listed a game "vs Brooklyn Nets", eight years before the Nets moved;
the 2008 standings put the Charlotte Hornets 23rd, a team that did not exist
that year; every all-seasons streak answer ended "franchises are named as they
are today", which was an honest description of a bug. ESPN keys a team by
franchise andteamsholds only today's names, so every row joined to it
read today's.The franchise table moved to a neutral module,
association.franchises,
because both packages need it -fetch/warehouse.pybuilds the
player_game_logview's abbreviations, and the query templates name teams
everywhere else - and CLAUDE.md keeps those two from depending on each other.
It gained each era's abbreviation (NJ, SEA, VAN), and two renderings of one
lookup:season_namefor Python andseason_name_sql, a SQL expression that
names each row for its OWN season, which a career log crossing a relocation
needs. A test checks the two against each other at every season either side
of every boundary. Both rename only whenteamsfiles today's name under
that id.Applied at every place a team name is printed: team game logs and quarter
scores, home/road standings, team metrics, record tables, all-seasons streaks,
player-matchup meeting logs, with/without stints (a stint across a rename names
both, "New Jersey Nets / Brooklyn Nets"),period_splitrows, and the
player_game_logview. The SQL agent's example queries inprompt.pyare left
alone, because changing them spends preamble budget.The view needs a warehouse reload to show it - it is built at load time -
and was reloaded with this change. Every site was perturbed back to today's
name and watched to fail; four of the first five came back MISSED until each
had a test of its own, and the view's test had to move from 2005 to 1997 to
catch anything, because the Grizzlies were already in Memphis by 2005. -
The router prompt's size is documented correctly, and budgeted.
query/router.pysaid the prompt was "~430 tokens" in its published
docstring and besideROUTER_NUM_CTX; it is 9,989 characters, about 2,500
tokens at the four characters a token the agent's budget is measured at,
against a 4,096-token window. NewROUTER_PROMPT_TOKEN_BUDGET(three
quarters of the window) and a test that fails when the prompt plus a long
question passes it - the router's counterpart toPreambleTooLarge, as a
test rather than a runtime check because the prompt is a constant. Not
measured with the model's tokenizer; the figure is an estimate. -
A stale comment in
router.route()no longer credits the 2.0 REPL with
theprevious_questionfollow-ups; it now says what arrives there
(Agent.last_question, None in both shipped callers) and who the branch is
for. No behavior change. Thedocs/usage.rstassists example now carries the
"(minimum 20 games)" qualifier the answer prints. The entry that filed it
said the answer also continued "Next: ..."; it does not, because the router
emitslimit1 for "who leads" (the web renderer's own note on why a
one-row ranking stays a sentence), and the template prints the list only
past one row. -
American spelling throughout
src/. 43 British spellings ("honour",
"behaviour", "labelled" and their forms) replaced there, and 24 more in the
tests, including the user-visible
check_scopetrace "cannot honour" and three published docstrings. The
router prompt and schema hash identically before and after, so no routing
moved. -
Two wrong comments corrected, no behavior change. The comment above
team_metrics.TURNOVERSsaid pre-2013team_season_stats.turnoverswas "the
player turnovers alone"; re-measured, it is the full count with team
turnovers in (equal to the boxtotalTurnoversseason sum for 24-27 of 30
teams, and to the player-only sum for none). The expression was right. The
player_season_stats_dedupedview andleaderboard.not_a_postseason_copy
now both give the dropped-row figure as 436 of 7,941 rows (340 of 7,845
player-seasons); one said 340 rows and the other 437. -
A team name is read for its season, and the question's own team beats one
the router could not ground. Found through "duren v nets 1h gameloh", which
fell through because the router wrote the opponent as "New Jersey Nets". That
was not new - the router wrote it identically in all seven replays since the
first - it was hidden, because the question used to be forced to the agent and
nothing read the slot untilperiod_splitdid.Two faults, and the second is the one that answered wrongly.
teamsholds
only today's 30 names, so every former name resolved to nothing - New Jersey
Nets, Seattle SuperSonics, Charlotte Bobcats, Vancouver Grizzlies, Washington
Bullets. And "Hornets" has belonged to two franchises, so matching today's
names answered "Hornets record 2008" with the 2008 Charlotte BOBCATS' 32-50;
the Hornets that season were New Orleans, and went 56-26. ESPN's ids belong to
the franchise, not the name - measured from where each team's home games were
played, id 17 is the Nets in New Jersey and Brooklyn alike - so the fix is a
list of the names the renamed and relocated franchises have carried and the
seasons they carried them (entities.FRANCHISE_ERAS), read for the season a
question is about. A name nobody held that season asks: "Hornets" in 2014 was
neither franchise. An id is only trusted where the warehouse files today's name
under it, because the first version renamed a test warehouse's Detroit Pistons,
filed under id 3, to the New Orleans Pelicans.A name whose city the router invented is read by its nickname, unless the city
contradicts it: "Portland Blazers" (ESPN writes Portland TRAIL Blazers)
resolves, "Los Angeles Kings" does not, because Los Angeles is two other teams.The precedence bug.
scope_from_questionread "nets" after "v" correctly
as the Brooklyn Nets, and then used it only ifopponentwas EMPTY - so a
router string resolving to nothing beat a team the question names outright.
The question now wins when the router's opponent is no team, or a team the
question never mentions. It also fixed a second shape nobody had diagnosed:
"andrew wiggins last 15 games vs warriors" arrived with the player and the team
SWAPPED between slots, and kept his name as the opponent.Fixing resolution exposed a third fault, caught by replaying before shipping.
"Keyonte George against blazers" arrived asteams=["Portland Blazers"], and
was answered correctly only BECAUSE that name failed to resolve. Resolving it
let the "slots already carry this team" rule - written forhead_to_head,
which readsteamsas its two sides - leave it there, andplayer_stat
answered his whole 54-game season instead of his 2 games against Portland.
With a player as the subject, a team after "vs" is his opponent wherever the
router filed it.Measured by replaying the last run's recorded router output through the fixed
entity stage and templates - exact for a change that lives entirely after the
router, and checked by reproducing all 261 rows of the recorded run on the
unchanged code first: 3 questions answer that fell through (Duren, Wiggins,
Duncan Robinson), and nothing else moves. -
period_split's average was inflated, and now counts the games he
played. As first shipped, a game only counted if he made a shot in that
period, so every scoreless quarter left the denominator: "RJ Barrett scored
250 points in the 4th quarter over 46 games, averaging 5.4" for a player who
played 57 and averaged 4.4. The total was right, which is why it read as
correct - it was graded correct in the replay that followed, and caught only
when the new per-game log printed his games and the count looked short. A
game with no shot data at all (2003's shots cover 986 of its games) is
excluded rather than counted as a confident zero, and a game he sat out is
no game.Two things the same replay found are fixed alongside it. A question asking
for a log - "rj barrett 4th qtr log", "vj edgecombe 1st quarter scoring by
game", 7 of the 11 the template answered - got a total and an average, and
now lists the games under a header that still answers the season. Andstat
is the router's one required slot, so it arrives filled on questions that name
no stat; "duren v nets 1h gameloh" came in withstat="none"and was refused
as asking for something other than points. Only a stat the question names is
kept, so "kd rebounds 4th quarter" is still refused rather than answered with
his points. -
A named player's quarter or half is answered, by a new
period_split
template. This was the largest content gap in the 261-query feed replay -
21 questions, every one forced to the agent because nothing answered the
shape. "rj barrett 4th qtr log", "victor wembanyama vs sacramento first half
log", "Devin Vassell nba player per game stats 1q". A TEAM's quarter has had
a template for a while (team_quarter_points, read from the official
linescore); a player's had none.Nothing was blocking it but a stale comment.
team_quarter_pointssaid a
player's quarter score "needs the plays-table LAG() derivation", and that is
not true:shot_chartalready carriesathlete_id,period,madeand the
shot's value, so the answer is a filtered sum.The value is read through
SHOT_VALUE_SQL, never guessed from the play's
prose, and that is the whole accuracy of the thing. Scored by looking for
"three point" in the description, per-period points match ESPN's own
linescores 76.8% of the time and the error is systematically -1: "makes
24-foot running jump shot" is a three that scores as two. Read off the shot's
own label and position it is 99.95%. Over a whole game that gap hides
inside a 98% figure; a quarter holds about ten field goals, so it does not.Validated a second way, which also settles a question the earlier entry had
to leave open. Summed over all periods INCLUDING overtime, a player's season
total matches his box score exactly for 550 of 578 player-seasons in
2026, mean error 0.138 points across a whole season - so the per-PLAYER
attribution inside a period is sound, not just the team-level total.Accuracy is a property of the season, and the template says so instead of
averaging it away. 2002 is refused, becauseSHOT_VALUE_SQLis NULL for
20,534 of its made shots and a sum over them means nothing (4.9%). 2016 is
refused too, at 76.5% - one quarter in four, the same season
reconstructed_boxsingles out. 2003-2006 and 2013 answer with the measured
figure attached. The other nineteen seasons run 99.2-100.0%.Points only. Rebounds and assists are not in
shot_chartat all, and
deriving them per period fromplayscarries its own fidelity per stat -
fouls reconstruct at 83% - so a question asking for them is refused with that
named as the reason rather than answered from a weaker source. A half is the
two quarters it holds and never overtime.The intent is assigned in
route()from the question's own words rather than
added toROUTER_PROMPTorROUTER_SCHEMA, because both are load-bearing on
every other question and a period is perfectly legible without the model's
help.CODE_ASSIGNED_INTENTSrecords that, so the reachability gate can tell
a deliberately unemittable intent from a dead one. -
A player against a team is answered instead of falling through. The
single biggest theme in the feed replay: 60 of the 186 questions that are not
answered correctly pit a player against a team, and three separate mechanisms
each half-handled it.player_matchupneeds two players. Given one and a team it had nothing to
answer with, and the rule that redirects those only ever read theplayers
LIST - the model routinely fills the singularplayerslot with an
opponentinstead ("keon ellis stats vs trailblazers", "Kd games vs
wizards", "De'angelo russell vs pistons"). Eight feed queries fell through to
the agent whereplayer_statandgame_loganswer them exactly, both
honouringopponent. The redirect is gated on the opponent being a team:
"jay huff game log vs Embiid" really is a matchup between two players, with
the second one in theopponentslot._is_team_namerequired a name's LAST word to be one of thirty nicknames, so
a team named any other way read as a player - "mathurin v det", "sam hauser
v mil", "pascal vs orlando". It now also accepts a city or an abbreviation,
matched against the WHOLE name and never the last word, because three real
players are surnamed Cleveland, Houston and Washington and a last-word rule
turns PJ Washington into a team.A misspelled team is deliberately still not matched, and that was
measured rather than assumed.difflibat 0.8 reaches the right team for the
feed's three typos ("taptors", "warriners", "blakers") - and also for 16 real
player surnames: Burks to Bucks, Hawkins to Hawks, Thornton to Toronto, Wheat
to Heat. The model puts bare surnames in that slot routinely, so those
collisions are live, and no cutoff separates them: "houstan"/"houston" and
"taptors"/"raptors" are both one edit in seven characters. Three queries is
not worth sixteen. -
The team abbreviations people actually write now resolve. ESPN's
teams
table abbreviates four teams "GS", "NO", "NY" and "SA", soGSW,NOP,
NYKandSASresolved to nothing and the team was lost from the
question entirely. The full "Los Angeles Clippers" did too - the router is
asked for full team names, and ESPN stores that one as "LA Clippers".An exact abbreviation also outranks a team whose name merely contains it.
ORLis Orlando's abbreviation and a substring of "New Orleans", and both
came back: an ambiguity offered over a question naming exactly one team, and
a chance to answer about the other. Two teams in one city stay ambiguous, as
they should - no team is abbreviated "LA". -
A narrowing the router has no slot for is refused instead of dropped.
check_scopecan only refuse a slot the router emits, andROUTER_SCHEMA
has no slot for a day of the week, a calendar holiday, an age, a minutes
condition or "since returning from injury" - so those words never reached it
and the template answered the un-narrowed question. Measured over 261 real
StatMuse feed queries, this was the single largest cause of a wrong answer:
14 of them, more than any other. "lebron james 2 3 pointers all-time vs jazz
on tuesdays" returned his career average against Utah over 48 games, with the
Tuesday, the threes and the "2" all silently gone; "anthony davis stats on
christmas" returned a whole season average; "most triple doubles before
turning 27" returned this season's leaders.A second pass covers the shapes the first measurement showed it had missed:
one game of a playoff series ("Ayton stats in game 4 playoff games" returned
his whole 10-game postseason) and a season named by ordinal ("his 18th
season", which the model read as the year 2018 and answered with that
season's league leaderboard). Both refuse. -
A calendar day without a year is now answered, not refused. "Desmond bane
march 17" returned his most recent game, dated 2026-04-12 - a month off, and
a different question. The year is not in the question and does not need to
be: a season fixes it. Season Y runs from October of Y-1 through June of
Y, so October to December belong toseason - 1and January onward to
season- this project's own numbering (current_season) applied to a
month. "Desmond bane march 17" resolves to 2026-03-17, andgame_log, which
honoursdate, answers "game on 2026-03-17, 16 PTS vs OKC".The first cut of this refused instead, on the reasoning that picking a year
the question never states is a guess. It is not: the season states it, and
refusing threw away an answer the warehouse holds. What genuinely cannot be
resolved still refuses, and the three cases are worth naming - a year the
question does state wins over the season's ("november 11 2019"); a date
that opens a window is a range, not a day, so "since January 31" is refused
rather than answered with one game; and a career question spans twenty
Octobers and fixes no year at all. February 31 is not a date either.They are read from the question text into the existing
situationslot,
never asked of the model - the same move_validate_sidemakes for the side
of the ball, and for the same reason: a new slot inROUTER_SCHEMAmoves
slots on unrelated questions, while a regex inroute()costs no prompt
tokens and cannot. No template listssituationinHONORED_SCOPING, so
each of these now refuses and falls through to the agent, which is the
ranking this project uses - a refusal beats a fluent wrong answer.Checked against 343 real questions (the 261-query feed plus the 83 routing
corpus cases): 14 feed queries match and no corpus case does, so nothing
that routes correctly today starts refusing. Separately,_AGENT_ONLYknew
q1but not1q, so "Duncan Robison 1q log" was answered with a whole-game
line; three more feed queries fixed by the mirror pattern. -
A triple-double abbreviation is no longer read as three-pointers. "luka
td3s home" answered with his points per game at home, becausetd3sbecame
shot_value: 3. The replay filed it under "condition dropped", which was the
wrong diagnosis - the venue was read and honoured correctly, and the fault is
the metric. Triple-doubles exist as a leaderboard metric, but nothing counts
them for one player and the season table they live on has no venue dimension,
so the question goes to the agent, which can derive them from box scores.
Spelled out, "triple double" already routed correctly; only the abbreviation
was unreadable. -
Splits, streaks and with/without read the rebuilt box line too - they were
the half that still called a rebuilt game a game he missed. Reading the
rebuilt lines landed for the per-game templates first; every template that
decides whether a player appeared kept askingplayer_box_statsfor
minutes, which a rebuilt row does not have. So the same season answered two
ways:game_loglisted Anthony Davis's 68 games of 2015 while
player_splitssaid he "was listed in 82 box scores in the 2015 regular
season but did not play in any of them". It now answers 68 games, 34 home
and 34 away.The teammate half was worse, because it answered fluently and no caveat
marked it.with_withoutasked the same question of the other player, so
every teammate in a rebuilt game read as absent and the game was filed on
the "without" side: "Anthony Davis without Eric Gordon, 2015" returned a log
padded with games Gordon played. It now returns 20 games, which is the
ground truth, and the split reads 68 played to 14 out.Three guards carried the fault, not one - the SQL that decides a player
appeared, the SQL that decides a game has no box score, and a Python filter
picking his games back out of a group - and each is now perturbation-tested
separately. A rebuilt row still has no minutes, so minutes are averaged over
the games that carry them rather than counting a rebuilt game as zero, and
the columns the rebuild gets wrong (UNGATED_ON_REBUILD: turnovers, fouls,
threes, attempts, rebound splits) are blanked rather than averaged in - a
split that reads 9 turnovers off a rebuilt line is the quiet version of this
bug, not a fix for it. A warehouse with noplayer_box_stats_filledview
falls back to the raw table and answers exactly as it did before. -
The power index stored 25 of ESPN's 90 rows a season, and nothing said so.
seasons/<s>/powerindexis a collection on ESPN's core API, which answers
{count: 90, pageIndex: 1, pageSize: 25, pageCount: 4, items: [...]}.
fetch_power_indexreaditemsfrom one response, so every season kept
ESPN's first page - 2024 held 25 rows covering 9 of 30 teams - and
team_outlookreported that as ESPN having no snapshot for most teams. A
short page and a short dataset are indistinguishable, which is how this
survived long enough forDATA.mdto record the missing rows as ESPN
"keeping only postseason teams". It holds all 30, in every snapshot.ESPNClient.get_collectionreads a collection to its end, asking for a large
page and then paging until it holds thecountthe response declares, and
warning if it never does.scripts/backfill_power_index.pyre-fetched
2017-2026: 250 rows to 630.get_jsonnow warns when it hands back an unexhausted page, naming the
URL and the page count. That is the part that generalizes: of the ten
endpoints this project reads, only this one is a paged collection - the other
two core-API calls are single resources, and the site and web APIs return
nested documents - so the guard exists for the eleventh, which will otherwise
look exactly as correct as this one did.team_outlookalso breaks a snapshot tie explicitly now. 2018's preseason and
regular-season snapshots are both stamped 2020-10-12, the day ESPN backfilled
them, one minute apart (07:47Z and 07:48Z) - so the regular season sorts last
today and no answer is wrong. Ordering by date alone rests the choice on that
minute; the tiebreak states it, and covers the case where the two stamps match
exactly. -
The 2000 playoffs are complete, recovered from a source the pull never
read. Games are discovered from each team's schedule, and ESPN's schedules
simply stop: the 2000 postseason ended on 2000-06-01, missing the whole
LAL-IND Final, WCF Games 6-7 and ECF Game 6. Its daily scoreboard is a
second, independent list and has those games, so a postseason pull now
makes a second discovery pass over it once the schedule's games are on disk,
scanning forward from the latest date stored. Nine games recovered
(70 -> 79), and every team in that postseason now matches ESPN's own season
totals exactly - the Lakers went from 15 games to 23.That the pass runs after the fetch is the whole fix, not a detail: the scan
needs a date to work forward from, and on a first pull the only dates that
exist are the ones the fetch just wrote. Run during discovery instead, it
found 70 ids for 2000 and none of the six Finals games - correct on a tree
that already held them, useless on a clean one.Postseason-only, and it costs a healthy season nothing: measured against
2024, the scoreboard and the schedules agree on all 82 games. 2001 recovers
only its Finals Game 5, because 23 days across that postseason's conference
finals and Final return no events at all - so that season is declared
postseason_partialand its answers now say what is missing rather than
stating a short series as fact. A 2001 regular-season question is unaffected;
that is a separate field for exactly that reason. -
A traded player's combined season row is rebuilt from his own stints when
ESPN's disagrees with them. ESPN's career endpoint returns one row per team
stint plus a combined row, and 19 of its 2,062 combined rows contradict the
stints they claim to combine - 13 are a byte-copy of a single stint, and 6
(1977-1983) are entirely NULL. Everything downstream preferred that row, so
it was the line every answer used: "Eric Murdock 1995-96 stats" read 9
games at 6.9 a game for a season he played 73 games of, and now reads
73 games and 647 points.player_season_statsis rewritten at load time, like the team-box repair, so
the deduped view and the leaderboard's own dedup both see it rather than one
being fixed and the other left reading the broken row. Every formula was
fitted against the whole warehouse before use - zero of 15,573 rows disagree
with any of them - and the two that could not be identified are refused
rather than approximated:avgMinutesis NULLed on a rebuilt row (it is the
one average with no season total behind it; a games-weighted mean is right
81% of the time and box-score minutes 61%), and a ratio over zero turnovers
is NULL rather than theinfDuckDB produces. -
Per-game answers now read the rebuilt box line, and say that they did.
Where ESPN serves an empty box score,player_game_logcarries the figures
rebuilt from play-by-play, andsingle_game_high,game_logand
threshold_countread them. "What was Anthony Davis's highest-scoring game
in 2015?" went from 0, to a refusal, to 43, on 2014-11-22 vs UTAH -
with the answer saying the figure is rebuilt rather than fetched. His 2015
game log lists 68 games where it used to report none, and "how many 20-point
games did he have that season" went from none, with a caveat to a real
count that says how many of those games were rebuilt.Counting is additive by construction and was checked rather than argued: an
empty line carries 0, so it can never clear a threshold of 1 or more. Over
2013-2018, across all seven readable stats, no athlete's count fell by a
single game and the league-wide totals rose (20+ point games, 15,978 to
18,488).Deliberately narrow, on measured grounds. Only the stats a rebuild gets right
are read (REBUILT_STATS): per player-game against the 22,646 games of 2015
whose real box score survived, free throws are exact to 0.0003, blocks 0.002,
rebounds and assists 0.004, field goals made 0.005, steals 0.009 and points
0.021 - but turnovers 0.080 and fouls 0.181, so those two are refused. Ask
for a player's fouls in an empty season and the answer says the lines exist
and were held back, rather than implying the data is missing.Season totals are not read anywhere, and that is the same measurement
seen from further away: a season is exact only when the net error over every
game is zero, so it lands right about half the time, and the error scales
with games played - a right total averages 31.6 games, a wrong one 55.6.
minutesis never invented; it prints blank on a rebuilt row, and the log
says so beneath the table. -
A row narrower than the rows after it no longer truncates the whole file.
storage.write_rowspassed its rows straight topa.Table.from_pylist, which
takes the Parquet schema from the FIRST row and silently drops every key only
later rows carry. Reproduced in isolation:[{"season": 2014}, {"season": 2016, "points": 299}]writes a file with nopointscolumn at all, while the
same two rows reversed keep it. Nothing raises, and the loss is permanent -
the value never reaches disk.It fired on real data because ESPN's career endpoint leaves a season out of
itstotalscategory when the player scored nothing: a career whose
earliest line is a scoreless one-game stint parses to a 26-key first row
followed by 51-key ones. Five files were written that way - Seth Curry's among
them, whose 2014 opens with a single game for Charlotte - each losing all 25
totals columns for that player's whole career, which is why those players were
missing from career and totals leaderboards entirely.Rows are now widened to the union of their keys before writing, and
homogeneous rows are returned untouched so the common path - about 218,000
files a pull, almost none of them ragged - pays nothing.This is older than the season-totals work and is not a regression from it:
measured against the 2026-09-11 warehouse, no row that had a total then is
NULL now, and 185 were repaired. The five files on disk are still truncated
until they are fetched again through the fixed writer. -
The warehouse answers from the rebuild where the stored line is empty.
New viewplayer_box_stats_filled:player_box_statswith the rebuilt
figures dropped into the 21,169 empty lines, under the stored table's own
column names so it is a drop-in, plus areconstructedflag marking exactly
those rows. Anthony Davis's 2015 - stored as 82 games of zeros - reads 68
games and 1,656 points through it, beside 14 rows correctly left as
did-not-play. ESPN's own season table says 68 and 1,656.Three things it will not do. It never substitutes into a real line, so a
player who genuinely scored 0 keeps his 0. It never inventsminutes, which
play-by-play cannot recover. And it drops the storedplusMinuson a
substituted row rather than passing it through: that column looks like
surviving data - it is not NULL, unlike every stat beside it - but across all
21,169 rows it takes exactly one value, 0, and every team-game sums to 0.0.
It is the same fabricated zero as the stats, wearing a different face.Neither view is in
KNOWN_TABLES, so the SQL agent reaches neither. A
substituted figure carries an obligation to say it was rebuilt, and an agent
writing its own SQL has nowhere to put that. -
2018's team box scores held their values under the wrong column names, and
now hold their own. Every non-empty 2018 row - both season types - was
shifted:assistsheld the game's blocks,stealsits turnovers,blocks
its fouls,foulsits flagrant fouls,fieldGoalPctits FT% and
freeThrowPctits 3P%. Per row, not on average:assistsequalled the
player-box block sum in 2,134 of 2,134 regular-season rows and the real
assist sum in 1. Separately, the teamturnoverscolumn is 0 in every row up
to 2012 andteamTurnoversholds a copy oftotalTurnoversthere rather
than the handful of turnovers charged to a team.Both are ESPN's, and neither is fixed by refetching - a clean pull reproduced
team_box_statsexactly, and the seasons either side of 2018 come from the
same code and the same column order with the right values. So the correction
is made at load time, in the newfetch/team_box_repair.py, from the game's
own player rows: a team's assists ARE the sum of its players' assists, and
ESPN agrees, its own team column equalling the player sum in 2,134 of 2,134
rows in 2017. The two percentages are recomputed from the made and attempted
columns beside them, which are right, rounded the way ESPN publishes them.What a user sees:
conditions._TEAM_LINEreadsAVG(t.assists), so every
2017-18 team split and with/without table reported the team's blocks as its
assists. Boston's home assists go from 5.2 a game to 23.7, the Lakers' from
4.8 to 24.4, Golden State's from 7.9 to 30.1; across the league, 4.8 to 23.0.Columns with no source are NULL rather than left holding another statistic:
the player box has no flagrant fouls, technicals, points in the paint or team
turnovers, so 2018'sflagrantFouls,technicalFouls,
totalTechnicalFouls,totalTurnoversandpointsInPaint(which is -1 in
every row) are cleared, as isteamTurnoversbefore 2013. A wrong value that
reads as a real one is the failure this project keeps producing, and nothing
insrcreads any of them.An empty team-game stays empty. Every Chicago and New Orleans game from 2013
to 2018 has an all-NULL team row beside player rows listing everyone as
having played with no minutes and every stat zero, and summing those would
turn "ESPN has no box score" into "this team recorded no assists". A row is
repaired only if it is non-empty AND its player rows carry minutes - fully
corrected or fully untouched, never half of each - which also leaves alone
the 117 all-NULL team rows whose player rows are real (Vancouver 1996,
Chicago 2000), a different fault with a different fix.scripts/check_team_box.pyverifies all of this against a built warehouse,
the waycheck_coverage.pydoes for the coverage floors. -
A box line rebuilt from play-by-play, for the games ESPN serves empty.
New viewplayer_box_stats_reconstructed, built at load time over the 1,024
events (of 1,025, all in 2013-2018) that have an empty box score and surviving
plays. No other ESPN source has these numbers: the CDN box score on a
different host serves the same zeros, the core API exposes no per-game athlete
statistics, and the athlete gamelog omits the games - which also shows the gap
follows the franchise, since Derrick Rose reads 0, 0, 0, 1, 61, 25 across
2013-2018 and Aaron Brooks 51, 65, 0, 1, 60, 26, each zero exactly in his
Chicago years.It is kept deliberately apart from
player_box_stats: its own view covering
only the empty games, snake_case columns so a derived value never looks like a
fetched one, and absent fromKNOWN_TABLESso the SQL agent can neither read
nor describe it. No template reads it. A player who appears in no play is
absent rather than zero, because a zero that reads as a real performance is
the exact bug this whole area is about.Accuracy is documented per column on the module, and is not uniform: free
throws 100%, blocks and rebounds 99.8%, assists 99.6%, field goals 99.5-99.6%,
points 98.3% per game - but season totals are exact only 51.7% of the time
(within 2, 72.2%), biased low, and 2016 is much the worst season. Minutes and
plus-minus cannot be recovered at all. -
A zero from an empty box score can no longer win a single-game high.
"What was Anthony Davis's highest-scoring game in 2015?" answered "0, on
2014-10-28 vs ORL" - fluent, dated, and false. Every Chicago and New Orleans
box score from 2013 to 2018 is stored with every
player listed as having played, no minutes, and every stat 0. Those lines
hold0rather than NULL, so they passed the "is not NULL" test beside them,
and where a whole team-season is empty the maximum over it is one of the
zeros.single_game_highnow reads only lines with minutes, the same line
_played()already drew inconditions- which is why streaks, splits and
with/without were never affected. Measured against the warehouse, no
unaffected answer moves: Davis's 2019 high is still 48, the 2015 league high
is still Kyrie Irving's 57, and Stephen Curry's 2015 high is still 51.The refusal it leaves says which fact is missing. "He has no games" is false
of a player who played 68 of them, so a season whose box scores are all empty
now answers "no 2015 regular season games with a box score in the
warehouse", and the note that follows gives the count and the years. -
Rows that are not games are no longer counted as games.
gamesholds
three kinds of row ESPN serves alongside the real ones, and each read path
filtered a different subset of them, so the same warehouse answered the same
question differently depending on which template got it. "How many times did
the Mavs play the 76ers in 2003" answered 3 for a season holding 2 - one game
stored under both230104006and400222658- and 1999-2000 matchups
counted 0-0 placeholders as meetings nobody won.There is now one filtered list,
real_games, built at load time
(fetch/real_games.py) and read byhead_to_head,conditions,
team_metrics, the team game log andteam_quarter_pointsalike. It drops
151 ofgames' 43,494 rows: 134 placeholders scored 0-0 with no winner, 23
team-slots naming an id no franchise has, 11 phantoms carrying a winner but
no box score and a date-only stamp, and the one same-day duplicate left after
those. Measured against the warehouse, the playoff records it corrects land
on the real ones - Orlando 1995 from 11-11 to 11-10, Seattle 1997 from 7-7 to
6-6 and 2000 from 2-6 to 2-3, Phoenix 1999 from 0-4 to 0-3, Portland 1999
from 8-6 to 7-6 - and Miami stops having a 1995 postseason it never played.
Chicago's 1999 game list falls from 100 rows to the 50 its standings line
says it played, and 2000's from 162 to 80.Two things it deliberately does not do. It does not drop a game that is
simply older than the box scores: every 1988-1992 game is stored date-only
with no box score, so the phantom rule fires only where that season and
season type have box scores at all. And it does not collapse season 1993,
which is a phantom SEASON rather than a phantom row - that stays with
coverage.pyand the cross-seasonQUALIFYinTEAM_GAMES_SQL, which are
the only things that can tell it from a real season.Needs a
data load: the list is built at load time, so a warehouse built
before this change does not have it. -
A season line served with no totals is repaired from a second endpoint.
246 rows inplayer_season_statscarriedavgPoints15.0 beside a NULL
points, so every career sum and totals leaderboard was silently short and
19 players - Seth Curry among them - were dropped from the career scoring
list outright, because every one of their season rows was NULL.The cause was two faults stacked, and only one is ESPN's. 53 of the 107
affected career files were served with only an averages category, and the
pull wrote that as-is and checkpointed it, so no later pull ever looked at
them again. The "a refetch does not fix it" recorded against this was sound
when it was measured - the surviving 2026-09-11 warehouse holds the same 246
NULLs - so what the two measurements together show is that ESPN's answer
CHANGED between 11 and 14 September. ESPN's own fault is the
smaller half: even a complete payload omits a line from itstotals
category when the player scored nothing (55 of the 62 remaining rows) or when
it is a traded player's combined row (the other 7).fetch_player_season_statsnow repairs such a line from
player_season_totals_url, the core per-season endpoint, before writing the
file. Fetched rather than derived, and the difference is measurable:
avg × gamesPlayedreproduces a known total exactly only 49% of the time
(14,464 of 29,534 rows, worst error 4), because theavg*columns are
rounded to one decimal. Seth Curry's career is the case for reading the real
number rather than trusting that spread: derived, it comes to 5,547.4999 -
right on the coin-flip, and 5,547 only becauseroundwent that way.The per-season endpoint has no team dimension, which is the whole
difficulty: it answers a traded player's combined figure against every one of
his stints, so David Wood's 21-, 4- and 37-game 1995-96 rows all come back
208 points.fill_missing_season_totalstherefore matches on games played
and fills exactly one row per season, refusing outright when two rows tie -
a stint that took the combined figure would read 208 points in 21 games with
nothing anywhere to say it was wrong.Measured against the warehouse: 110 requests repair 226 of the 246 rows
(123 from the career endpoint alone, which now serves totals for 87 of the
107 files, and 103 from the per-season one). The remaining 20 are 14
one-game lines that scored nothing and the 6 all-NULL combined rows from
1977-1983, which ESPN answers 404 for. Career-leaderboard drops go from 19
players to 0. Backfilled withscripts/backfill_season_totals.py. -
A single-game high keeps the player the question named. "most points
curry scored in a game this season" came back from the router as
single_game_highwith no player slot at all, and the answer was the
league's high - Bam Adebayo's - to a question about one man. The player is
optional for that template (an empty slot means the league), so nothing
downstream restored it, andplayers_named_incould not: "curry" is six
players and it refuses to guess between them.The subject is now read from the question's grammar - a name before a scoring
verb, or carrying a possessive - and handed to normal resolution, which asks
"did you mean Seth Curry or Stephen Curry?". A word scan could not do this:
"best" is Travis Best, "game" is Jaron Blossomgame, "high" is Haywood
Highsmith and "single" is four players, so scanning would answer "the highest
scoring game by a player this year" about somebody. Read from the text, so
ROUTER_PROMPTandROUTER_SCHEMAare unchanged. -
Per-game leaderboards for points, rebounds, assists, steals and blocks
apply a games minimum. These five ranked every board unqualified, so the
fewest games was the easiest route to the top of one: "who led the league in
rebounding in 2001" answered Danny Fortson, who played 6 games, where
Dikembe Mutombo led it over 79, and "who led the 2023 playoffs in scoring"
answered Kawhi Leonard on 2 games rather than Devin Booker on 17. Measured
over 1994-2026, two regular-season boards and 15 postseason ones were led
from under the floors. They now use the same 20 games, and 5 in the
postseason, that every newer per-game metric already used, and the answer
names the qualifier it applied ("minimum 20 games") the way a percentage
already named its attempts.Counts are deliberately left alone. A season total, a double-double count
and a cumulative NetPoints figure need volume to rank at all, and no board
of one was ever led from under these floors - the two triple-double boards
that were are right, since nobody records more triple-doubles than he plays
games. Career leaderboards already qualified and are unchanged.The five are now built by the same helper the newer per-game metrics use, so
"a per-game metric" has one definition rather than two that can disagree,
and a test over the whole registry fails if another arrives without both
floors. That is what was missing: the omission read as deliberate, because a
metric naming a qualifying column and no threshold is how this registry says
"rank this unqualified". -
Fix: "without X and Y" was answered about X alone. The router read only
the FIRST name out of a "without" or "with" phrase, so "Celtics record
without Tatum and Brown" arrived aswithout='Tatum'and "Lakers record
without Lebron and AD this season" as'Lebron'. The answer then covered the
games one of the named players missed and said nothing about the other: a
narrower question, answered fluently, with the dropped name nowhere on the
page. Real questions have this shape - the StatMuse feed has "hornets record
when brandon miller and lamelo and knueppel play this year".router._names_afternow reads every name the phrase holds, joined by "and",
"or", "nor" or a comma, and the three templates that honourwithout
(with_without,game_log,player_stat) require all of them: a game
counts as "without" only where NONE of the named players played, and
with_without's "with" row only where every one of them did. The games in
between - one played, one sat - go on the other row, which is what stops a
two-player question being answered about one player.with_withoutcounts
only the time the named players were all on the same team, as it already did
for one, and its rows name which side is which ("Tatum and Brown out"
against "Tatum or Brown played").Read out of the question text, so
ROUTER_PROMPTandROUTER_SCHEMAare
unchanged (both hashed before and after) and no other question's routing can
have moved - the same lever_validate_seasonand_validate_sideuse. The
phrase parser also learned the question words ("how", "what", "who"...) as
name terminators, so "without Tatum and how many wins" still reads one name
rather than making a teammate out of the tail of the sentence.The
withoutandwith_playerslots are lists of names now; a bare string
is still read as one name (entities.teammate_names), since slot values are
advisory everywhere else here.player_statandgame_logreport
data["without"]as a list, andwith_withoutgaineddata["teammates"].
scripts/check_routing.pygained the two-name case, and theISSUES.mdP1
entry it came from is closed. -
Docs: three stale comments describing a gap
period_splitalready
closed.router.py(the_HALF_WORDScomment and the one above
_is_team_quarter_points) and theteam_quarter_pointsdocstring in
templates.pystill said a player's quarter or half had no template and
would need a fragile plays-tableLAG()derivation - true when written, and
false since31fe043addedperiod_split, which readsshot_chart
instead. Comments and docstrings only; no behavior changed.