Skip to content

mlb_groups

Choose a tag to compare

@saiemgilani saiemgilani released this 27 Sep 03:30
· 17 commits to main since this release
f64d5c7

MLB conference, division and subdivision reference.

126 seasons (1901–2026) · 260 assets · 1.1 MB · last asset written 2026-09-28

Upstream source: the league's membership source and every source's names and ids for its groups (see the build repo).

Season key. the single calendar year

Note. A reference table, not observations. Four tables: groups (one row per SDV group lineage, {league}:{slug}), group_seasons (each group's name, abbreviation and parent as of that season), group_aliases (every source's ids and names for a group, with validity windows) and team_group_seasons (each team's subdivision, conference and division by season, one file per season). Historical names come from dated, cited rows: ESPN, stats.ncaa.org and CFBD all show today's names for past seasons. Schema: sdv-reference-data CONTRACT.md.

Files in this release

Format Files Name pattern
csv 130 {dataset}_{season}.csv
parquet 130 {dataset}_{season}.parquet

This tag carries 4 distinct datasets under one release: mlb_team_group_seasons, mlb_groups, mlb_group_aliases, mlb_group_seasons.

Direct download (no auth, no API key):

https://github.com/sportsdataverse/sportsdataverse-data/releases/download/mlb_groups/mlb_team_group_seasons_2026.parquet

How to load it

Surface Call
Python from sportsdataverse.mlb import load_mlb_groups
load_mlb_groups()
Python from sportsdataverse.mlb import load_mlb_group_seasons
load_mlb_group_seasons()
Python from sportsdataverse.mlb import load_mlb_group_aliases
load_mlb_group_aliases()
Python from sportsdataverse.mlb import load_mlb_team_group_seasons
load_mlb_team_group_seasons(seasons=[2026])
R baseballr::load_mlb_groups()
Any pd.read_parquet("https://github.com/sportsdataverse/sportsdataverse-data/releases/download/mlb_groups/...") — the assets are plain files on a public URL

How it is produced

  1. Capture — sdv-reference-data: fetch() in sdv_reference/leagues/mlb.py snapshots the sources into raw/mlb/.
  2. Build — sdv-reference-data: scripts/pipeline/20_build_tables.sh mlb (offline, from raw/ and cited curated/ rows).
  3. Publish — scripts/pipeline/30_publish_releases.sh mlb in sdv-reference-data uploads the assets to this tag.

The code that writes these files:

Automation

No scheduled automation is attached to this tag: it is refreshed on demand.

What depends on it


Part of SportsDataverse. Every release in this repository is a public, versioned data asset: the files are stable URLs you can read straight from R, Python, or anything that speaks HTTP. Issues with the data belong on the producing repository listed above; issues with a loader belong on that package's repository.