-
Delete:
data/players.csvdata/teams.csvdata/matches/replay/*data/matches/replay_failures.csv
-
Update:
config.R(league ID, starting patch, period, any tier 1 or 2 events OpenDota marks as excluded, andperiod_names/period_series/period_median_series, which describe how many series a team plays in each period and come from the published format rather than from the data)data/heroes.txt(extractnpc_heroes.txtfromdota 2 beta/game/dota/pak01_dir.vpk)data/stats.csv(with the new points system)data/prefixes.csvdata/suffixes.csvdata/banners.csv(with the emblem colours and order of each War Banner)data/qualities.csv(the quality ladder, in order, with the roll weights frommethodology.mdsection 13)data/traits.csv(each trait's own bonus and what it does to its neighbours)data/rolls.csv(every roll operation the game can offer)
-
Update as the tournament progresses:
config.R(current_periodandteams_eliminatedbetween periods)
Only the three files above are read from disk when they exist, so they are the only ones a new tournament needs deleted. The rosters and teams are kept because they can be corrected by hand; the parsed replays because they take hours to rebuild. Everything else is refetched and overwritten on every run.
data/observations/ is not part of any of this. It holds measurements taken by
hand to settle a modelling question — currently emblem_quality_rolls.csv, the
recorded rolls behind the quality weights. Nothing reads it at run time, and it
describes the game rather than the tournament, so a new event leaves it alone.
The match list and the match data both come from OpenDota's SQL explorer, a few
whole-set queries rather than one request per match, cheap enough to repeat every
run. A patch is a span of time, so starting_patch is turned into a date using
the release dates OpenDota publishes. The explorer stores a replay's cluster and
salt rather than the url built from them.
Two things to know when supplying the data above:
data/heroes.txtmust parse withvdf, which only handles one closing brace per line. If it fails, look for a line holding two and split them.- Each prefix and suffix needs a matching condition in
src/compile-match-data.R. Renaming or adding one without that fails the run rather than silently producing an empty column.the Cruelis currently unimplemented. - Each trait needs one in
docs/calc.jsfor the same reason, since only the numbers live indata/traits.csvand the conditions cannot. Row order indata/qualities.csvis the ladder emblems improve along, so it matters.
- Take a pull from
https://github.com/skadistats/clarity-examplesand merge it intohttps://github.com/VirenDias/clarity-dota2-fantasy. - Update for any new fantasy developments and fix any issues.
- Compile the
.jarand add it toutils/fantasy.jar.
The parser is the only source of stat values, so a property renamed by a new patch produces a silently empty column rather than an error. After the first matches are parsed, compare a few against the same fields on OpenDota before trusting a run.
Run fantasy-stats.R. It fetches whatever is missing, which on a first run means
downloading and parsing every replay, and writes:
data/match_data.csv— one row per player per match, with the points each emblem stat scored and which prefixes and suffixes appliedresults/role_stats.csv— recency-weighted average and standard deviation per team, role and metric, with the number of matches behind each figuredocs/data.json— the same matches as unaggregated role-games, for the calculator
A replay is downloaded once and kept only as its parsed csv, so a rerun resumes wherever the last one stopped. Matches whose replay cannot be downloaded or parsed are reported and dropped, but the run stops if fewer than 95% survive, since that points at the parser rather than at the matches.
Replays that keep failing are counted in data/matches/replay_failures.csv and
left alone after a few runs, sooner for one Valve no longer serves than for a
transfer that stops early, which resumes where it stopped. Delete that file to
try them all again. Matches OpenDota has no data for are only reported, never
recorded, since asking again costs nothing and OpenDota fills them in late.
docs/ is a static page with two tools over one shared War Banner. Serve it and
open it:
python -m http.server 8000 --directory docs
Set each banner slot to the emblem you rolled — its stat, quality and trait. The
multiplier is derived, including what neighbouring emblems do to it, and shown so
it can be checked against the game. Everything recalculates in the page from
data.json, so nothing needs rerunning in R.
Reroll ranks every roll operation the game can offer, against every banner, by the expected change in points. Rolls are compared across all three banners because tokens are shared and a point is a point wherever it lands. Mark the three you were offered and it names the one to take, or says to take none. Every row opens to show all of its outcomes with their odds, so the numbers can be checked by hand.
Rosters ranks every combination of one team per role, with the prefix and suffix that score best for it.
A score is the best two matches of a series, then the best series of the period,
per role. Both maxima are evaluated exactly rather than sampled: a series score
is always the sum of two of a team's games, so every outcome is one of at most
n(n+1)/2 atoms with a closed-form probability. Games are drawn weighted by
recency, which weights their stats and their prefix and suffix triggers together.
There is one column per number of series a team might play, since that is the one
input the tournament decides rather than the data. The best pair is chosen at the
median and held across the others, because you commit to one title rather than
one per scenario. Read methodology.md before trusting the suffix column —
the Clutch is knowingly overstated there.
An indicator with no data at all is dropped from the export on the evidence
rather than by name, so the Cruel reappears on its own once the parser supplies
it.
The Compiled Data tab renders as markdown Reddit will accept:
node src/reddit-tables.js
It writes results/reddit-tables.md, or add --print to send it to stdout as
well. No dependencies and no build step — it reads docs/data.json and calls
docs/compiled.js, the same code the tab runs, so a posted table cannot drift
from the site. Paths resolve against the script, so it runs from any directory.
The two per-Team tables put the teams across the columns by tag, and the Emblem
Stats and Titles down the rows — one table per Role rather than one per Role and
Emblem Colour. That is the wider of the two arrangements, but it collapses
thirteen tables into six and stops repeating sixteen team names in every one.
Matches per Team and Role carries both the name and the tag, so it doubles as
the key to everything below it.
The table titles are read out of docs/index.html rather than kept a second
time, and the run stops if it cannot find all five tables there. Only headings
and tables are written — a post carries its own framing.
Every asset in docs/ is served with a ?v= stamp, so a browser holding a
cached index.html loads one consistent set rather than pairing new data with
last week's code. The stamp is an md5 digest of calc.js, rolls.js, app.js,
style.css and data.json, written into index.html by stamp_assets() at the
end of the export. Being a digest, it moves when any of those files change and
holds still when they do not.
It is only recomputed when the export runs. Editing anything in docs/ and
pushing without a run therefore ships new code under the old stamp, and anyone
who has already loaded the page keeps the old copy from an unchanged URL. So
after a change to the page, run the export before pushing:
source("fantasy-stats.R")Only docs/index.html changes if the data has not moved. A pre-commit hook
calling stamp_assets(asset_version()) would remove the need to remember, but
there is deliberately no hook here.
MIT — see LICENSE. Use it, fork it, host it; just keep the copyright notice.
If you build something on this, a link back is appreciated.
The licence covers the code and the written analysis. It does not cover:
utils/fantasy.jar, which bundles clarity, clarity-protobuf, snappy-java, fastutil and slf4j, each under its own licencedata/heroes.txt, extracted from the game files and owned by Valve- the match data, derived from OpenDota and from Valve's replays