Skip to content

baseballr 2.0.0

Latest

Choose a tag to compare

@saiemgilani saiemgilani released this 27 Aug 04:12

New features

  • statcast_search() no longer renames Savant CSV columns by position.
    Savant's own header row is trusted as-is (its names are already the
    canonical identifiers), so a column inserted mid-export can never again
    silently shift every downstream value — the "n_thruorder_pitcher showing
    AGE" failure class (#337, #354, #371, #390, #408, #416). Genuinely new
    Savant columns now arrive under their own names with an informational
    message.
  • statcast_search_minors() and statcast_search_wbc() — the minor-league
    and World Baseball Classic Statcast searches (separate Savant routes, same
    columns) (#339, #395).
  • statcast_pitch_colors() — Baseball Savant's pitch-type color palette as
    a tibble for consistent pitch charts (#356).
  • Eleven load_mlb_*() loaders for the pre-computed MLB model datasets on
    the sportsdataverse-data releases (expected stats/HR, batter projections,
    Stuff+/Command+/xERA, OAA, catcher framing, RE24 matrix, WE table, WPA),
    mirroring sportsdataverse-py's loader surface.
  • mlb_stats() and mlb_teams_stats() gain sit_codes for situational
    splits via stat_type = "statSplits" (vs L/R, home/away, RISP, ...)
    (#278, #304, #383). Note the upstream leaders endpoints ignore sitCodes.
  • fg_projections() — FanGraphs player projections (Steamer, ZiPS, ATC,
    THE BAT and rest-of-season variants; batting and pitching) (#182).
  • ggpitchzone() — pitch-location plot from the catcher's perspective with
    the strike zone overlaid, colored with the Savant pitch-type palette
    (#260).

Bug fixes

  • mlb_pbp() no longer back-fills the at-bat-level matchup.postOn* /
    matchup.splits.menOnBase columns across at-bat and half-inning
    boundaries — bases-empty plays previously inherited baserunner ids from
    the next at-bat (#263). This corrects long-standing output.

  • mlb_pbp()'s count.balls.start / count.strikes.start are now true
    pre-pitch counts (the previous event's post-pitch count within the
    at-bat, 0-0 before the first pitch). They previously carried the
    post-pitch counts under a pre-pitch name (#131, #252).

  • mlb_game_info() degrades gracefully for games that have not started:
    venue/weather/status return with NA attendance and first-pitch fields
    instead of erroring (#363).

  • mlb_pbp() works for pre-2010 games again: play events are tied to their
    at-bat by position instead of an endTime join the older feeds don't
    carry (#233). Modern-game output is unchanged.

  • mlb_pbp(add_base_state = TRUE) appends per-event pre-pitch base-occupancy
    columns (pre_on_1b / pre_on_2b / pre_on_3b, runner MLBAM ids)
    reconstructed from the feed's runner-movement records, including the
    extra-innings automatic runner. Validated pitch-for-pitch against
    Statcast's own on_1b/2b/3b columns: 4,388/4,388 states agree across a
    full 15-game slate (#276).

  • bref_standings_on_date() supports every divisional era: pre-1969
    league-only standings, 1969-1993 East/West, and 1994+ East/Central/West.
    Tables are named from the page's own section headings, and requesting a
    division that did not exist for the date errors with the era's available
    options (#253).

  • mlb_rosters() documents the mlb_teams() team-id lookup (#376).

  • Added read-only Fox Sports "Bifrost" MLB wrappers (fox_mlb_*()) over api.foxsports.com/bifrost/v1/mlb/*: fox_mlb_team_roster(), fox_mlb_team_stats(), fox_mlb_team_gamelog(), fox_mlb_standings(), fox_mlb_league_leaders(), and fox_mlb_odds(). They flatten Fox's layout-oriented JSON (sections → tables → rows → cells) into tidy baseballr_data tibbles. Fox does not expose MLB play-by-play or boxscore via event/{id}/data, so those are intentionally omitted. Parallels the cfbfastR / hoopR / fastRhockey / sportsdataverse-py fox_* families; reverse-engineering notes + an OpenAPI 3.1 spec live in the sdv-internal-refs repo.

  • Added a full ESPN MLB wrapper family (espn_mlb_*(), 100+ functions) mirroring the naming and structure of the sister SportsDataverse packages (hoopR espn_nba_*/espn_mbb_*, wehoop espn_wnba_*/espn_wbb_*, cfbfastR espn_cfb_*). It covers ESPN's three public hosts:

    • Game data (site.api.espn.com): espn_mlb_scoreboard(), espn_mlb_pbp() (pitch/at-bat level), espn_mlb_team_box(), espn_mlb_player_box(), espn_mlb_game_all(), espn_mlb_game_rosters(), espn_mlb_teams(), espn_mlb_team_current_roster(), espn_mlb_standings(), espn_mlb_betting(). Baseball-specific extractors with no basketball analogue: espn_mlb_game_probables() (probable / announced starting pitchers) and espn_mlb_game_info() (venue, attendance, game duration, and the umpire crew).
    • Reference data (sports.core.api.espn.com): athletes, coaches, seasons, franchises, draft, futures, leaders, positions, powerindex, events, plays, per-competitor game detail, and the season_* family.
    • Athlete data (site.web.api.espn.com): espn_mlb_player_overview(), espn_mlb_player_gamelog(), espn_mlb_player_splits(), espn_mlb_player_stats(), espn_mlb_team_stats().
      Box scores are returned as wide, clean-named baseballr_data tibbles (one row per team for espn_mlb_team_box(); one row per athlete-side, tagged stat_group, for espn_mlb_player_box()), with the baseball batting/pitching/fielding stat groups parsed natively. ESPN requests share an internal httr2 layer (.retry_request()) that honours options(baseballr.proxy = ...). Live tests are gated behind ESPN_MLB_TESTS=1 (skip_espn_test()).
  • Added a full ESPN College Baseball wrapper family (espn_college_baseball_*(), 70 functions) covering ESPN's NCAA college-baseball endpoints (sport baseball, league college-baseball). They are thin twins of the espn_mlb_*() family over the same league-parameterized helpers, so return shapes match their MLB counterparts. Coverage includes game data (espn_college_baseball_scoreboard(), espn_college_baseball_pbp(), espn_college_baseball_team_box(), espn_college_baseball_player_box(), espn_college_baseball_game_all(), espn_college_baseball_game_rosters()), teams / standings / rankings / news / conferences, the team-detail and athlete families, the core-v2 reference graph (seasons, season types, groups, venues, coaches, calendar), per-competitor game sub-resources, and tournaments (the College World Series is an ESPN tournament). Endpoints ESPN does not serve for college baseball (injuries, win-probability, betting futures/odds, athlete contracts, transactions, draft, depth charts) are intentionally omitted. A new helper most_recent_college_baseball_season() supplies the season default. Endpoint/return catalog and captured sample bodies live in the sdv-internal-refs repo.

Bug fixes

  • bref_standings_on_date(), bref_daily_batter(), bref_daily_pitcher(), and bref_team_results() now retry on HTTP 429 (Too Many Requests) from Baseball Reference with exponential backoff (up to 3 tries) via a new bref_read_html() helper that wraps httr2::req_retry(). Previously, a 429 was swallowed by the tryCatch into a misleading "Invalid arguments or no data available" message, leaving the returned data as NULL (#344). Error handlers also now print the original error message for diagnosis.
  • The NCAA family (ncaa_schedule_info(), ncaa_roster(), ncaa_teams(), ncaa_team_player_stats(), ncaa_pbp(), ncaa_game_logs(), ncaa_lineups(), ncaa_park_factor()) can fetch stats.ncaa.org again — clearing the new Akamai 403 / soft bm-verify block — after the site simultaneously (a) tightened its Akamai bot protection and (b) redesigned its pages (#410). Each affected scraper was also updated for the new page layout (details below); clearing the block needs the optional chromote + Google Chrome browser fallback.
    • Akamai bypass. stats.ncaa.org now blocks every httr2/curl request — a hard 403 or a soft HTTP-200 bm-verify interstitial — on the client's TLS/sensor fingerprint, so no header set can clear it, and a vanilla headless Chrome is blocked too (Akamai fingerprints the HeadlessChrome user-agent and the navigator.webdriver flag). request_with_proxy() now falls back to a stealth headless-Chrome fetch (real Chrome user-agent, navigator.webdriver hidden) via the optional chromote package, wrapping the rendered HTML in a synthetic response so every caller is unaffected. The browser session is cached and reused. chromote + Google Chrome are an optional (Suggests) dependency; when absent the scrapers emit a clear install message. The fast httr2 path is used whenever Akamai allows it and resumes automatically if the edge relaxes.
    • Page-redesign parsers. Each scraper was updated for the new layout: ncaa_schedule_info() selects the schedule table by its header columns (it moved out of <fieldset> into a plain <table>); ncaa_team_player_stats() reads the migrated /teams/{season_team_id}/season_to_date_stats #stat_grid; ncaa_pbp() follows the "Play By Play" tab and parses the table.table inning tables; ncaa_game_logs() reads the /players/{id} #game_log / #career_totals grids; and ncaa_lineups() derives the batting order from the individual_stats box scores.
  • ncaa_game_logs() and ncaa_team_player_stats() gain type = "fielding" support (previously only "batting" / "pitching").
  • edge_frequency(df, group = ...) now groups by the column named in the group argument. It previously grouped by a literal column named group (.data$group), so passing a group (e.g. "pitcher") errored unless the data happened to have a group column and otherwise ignored the argument.
  • statcast_search() now assigns Baseball Savant's columns length-tolerantly and recognizes the newest Savant column definitions, so columns added to (or reordered in) the CSV export no longer break the function with a "can't assign N names to an M column data.table" error (#337, #354, #371, #390).
  • statcast_search() now recognizes Baseball Savant's miss_distance column (the bat-tracking swing-and-miss distance). Savant inserted it mid-frame — between swing_length and estimated_slg_using_speedangle, not as a trailing column — so the length-tolerant positional rename labeled it (and every column after it) one position off: miss_distance dropped out and the trailing bat-tracking columns were silently mislabeled. miss_distance is now in the canonical column list at its real position, so it comes through under its own name and the columns after it line up again (#408).
  • statcast_search() (via process_statcast_payload()) now normalizes blank character fields to NA. Baseball Savant exports empty fields as empty strings, so character columns such as events, des, and description were coming back as "" rather than NA (numeric columns were unaffected). This broke downstream helpers that treat missing categorical values as NA -- e.g. metrics_linear_weights_savant() was not folding balls/strikes into the events column (#275).
  • statcast_search() (via process_statcast_payload()) returns data again. Baseball Savant's CSV export ships the pitcher and fielder_2 columns twice (the trailing repeats carry identical values), and after the positional header rename those duplicate names survived into the payload. dplyr::across()/mutate() reject a frame with duplicate names (Can't transform a data frame with duplicate names), so the function now drops the duplicate-named columns (keeping the first occurrence) before any column-wise transform, class-aware for the data.table payload returned by csv_from_url().
  • sptrc_team_active_payroll() and sptrc_league_payrolls() updated for Spotrac's new /payroll/_/year/<year>/ URLs and changed table schema; parsing is now resilient to column-order changes and both functions return data again (#392).
  • Wrappers that build their result inside tryCatch() now initialize the return value first, so an API error returns an empty value with a cli message instead of an object '<var>' not found error.
  • mlb_game_timecodes() no longer returns NULL; it renamed its single column by the literal name ".", which newer R versions no longer use for the coerced column, so it now renames by position.
  • mlb_homerun_derby() and mlb_homerun_derby_bracket() return data again; their column selections are now wrapped in dplyr::any_of() so the dropped num_batters column (no longer in the MLB Stats API response) no longer errors the parse.
  • The NCAA functions (ncaa_*()) work again. stats.ncaa.org moved behind Akamai, which returns HTTP 403 "Access Denied" to requests lacking modern browser headers; .ncaa_headers() now sends the current client-hint / fetch-metadata headers (sec-ch-ua, Sec-Fetch-*, Upgrade-Insecure-Requests), and ncaa_teams() / ncaa_team_player_stats() now request over https. request_with_proxy() now sleeps 5 seconds after every request (the NCAA edge aggressively rate-limits / IP-bans) and accepts a proxy argument (a URL string or a list of httr2::req_proxy() args, defaulting to getOption("baseballr.proxy")) to route and rotate requests through a proxy. (Verified: ncaa_roster(), ncaa_teams().)
  • ncaa_teams() no longer returns NA ids. stats.ncaa.org migrated team links from the legacy /team/{team_id}/{season_id} pair to a single modern /teams/{season_team_id} resource, so the previous str_extract("(\\d+)/(\\d+)") matched nothing and team_id/season_id came back NA. The function now extracts ids from whichever url shape is present (legacy rows populate team_id/season_id; modern rows populate the new season_team_id column) so the id columns are never all-NA when the site flips formats.
  • ncaa_team_player_stats() builds its request correctly and fails gracefully (#379). The batting branch previously passed id twice and never sent year_stat_category_id, so it could not return batting stats; both branches now key /team/{team_id}/stats on the season id and the per-type stat-category id (batting_id / pitching_id) from the season-id lookup. The stats endpoint is currently gated behind Akamai's bm-verify interstitial challenge (verified: the URL and ids are correct and work in a real browser, but a static request — even with valid session cookies carried from the roster page — is served the challenge, not the data). When challenged, the function now emits an informative cli warning and returns an empty tibble (via the new internal .ncaa_is_interstitial() guard) instead of silently scraping a table-less page.
  • fg_batter_leaders() returns data for FanGraphs' handedness-split leaderboards (month = 13 for vs LHP, month = 14 for vs RHP) instead of NULL (#323). The split boards return a much narrower column projection (~83-97 columns) than the full board (~475), and the function's rename()/select() referenced columns absent from that projection, so the parse errored and was swallowed to NULL. The leading rename and column selection now use dplyr::any_of(), so a column missing from a narrower split is skipped rather than failing the whole parse. The same any_of() hardening was applied to fg_pitcher_leaders() and fg_fielder_leaders(), which shared the strict select() / rename() pattern (verified live: fg_pitcher_leaders(month = "13") returns data).
  • fg_batter_leaders(), fg_pitcher_leaders(), fg_fielder_leaders(), fg_team_batter(), fg_team_pitcher(), and fg_team_fielder() now honor startdate/enddate out of the box (#326). FanGraphs only applies a custom date range when month = "1000", so supplying dates with the default month = "0" silently returned the full-season board. month is now set to "1000" automatically when a date range is requested without an explicit month, and the month documentation covers the special values ("13" vs LHP, "14" vs RHP, "1000" date range). Verified live: a 2023-03-26..2023-04-26 pull returns the date-ranged board (max 118 PA) instead of the full season (max 753 PA).
  • ncaa_park_factor() no longer errors mid-pipeline when the NCAA schedule is unavailable (#302). ncaa_schedule_info() can come back empty -- stats.ncaa.org rate-limits / blocks automated requests and its schedule schema has drifted -- which left the schedule frame without the home_team_id / score columns the function mutates on, throwing a confusing Column 'home_team_id' not found (historically 'opponent' not found). The function now guards the schedule frame and returns an empty result with an informative cli warning instead, in line with the rest of the package.
  • fg_guts(), fg_park(), and fg_park_hand() return data again. FanGraphs replaced the legacy ASP.NET grid (HTML id GutsBoard1_dg1_ctl00) with a modern .table-scroll data grid, so the hard-coded id selector matched nothing; they now target the current grid, use https://, and fetch through the okhttp User-Agent (below) so the Cloudflare challenge doesn't block them.
  • All FanGraphs functions return data again under FanGraphs' current Cloudflare posture. As of 2026-06-03 FanGraphs serves a Cloudflare JS challenge (HTTP 403, cf-mitigated: challenge) to every unrecognized client — including the plain/library User-Agent that previously worked — and no header/TLS tweak passes it (it requires a JS runtime). The challenge exempts the okhttp client the FanGraphs mobile app uses, so the FanGraphs requests now send User-Agent: okhttp/4.12.0 and get HTTP 200 with normal JSON. A dedicated internal fg_api_call() helper (okhttp UA) carries the leaders/team functions (fg_batter_leaders(), fg_pitcher_leaders(), fg_fielder_leaders(), fg_team_batter(), fg_team_pitcher(), fg_team_fielder()); the game-log functions (fg_batter_game_logs(), fg_pitcher_game_logs(), fg_milb_batter_game_logs(), fg_milb_pitcher_game_logs()) and the .aspx scrapers send the same UA. The shared mlb_api_call() keeps its plain package UA for the MLB Stats API (which ignores it). (#404, #402, #397, #389, #385, #384, #373, #369, #361, #358, #353, #349, #343; okhttp-UA approach from PR #405.)

Documentation

  • Added a description column to the @return tables of all currently-returning functions (matching the wehoop/hoopR documentation style), regenerated from live API responses so columns, types, and descriptions reflect the current output.

Internal / infrastructure

  • Migrated the entire HTTP layer from httr to httr2 and dropped httr as a dependency. httr::modify_url() -> httr2::url_modify_query(); the MLB/FanGraphs helpers and the NCAA helpers (request_with_proxy() / check_status()) now use httr2::request() / req_retry() / req_perform() / resp_body_string() / resp_status(). request_with_proxy() gained a proxy argument (a URL string or a list of [httr2::req_proxy()] args, defaulting to getOption("baseballr.proxy")) for routing NCAA requests through a proxy.
  • Migrated the package's internal code from the magrittr pipe (%>%) to the native pipe (|>); Depends is now R (>= 4.1.0).
  • User-facing messages migrated to the cli package.
  • Column selections that drop known-transient columns now use dplyr::any_of() for resilience to upstream schema drift.
  • Added project documentation and community health files (CLAUDE.md, CONTRIBUTING.md, CODE_OF_CONDUCT.md, .github/copilot-instructions.md, pull request template), refreshed the _pkgdown.yml reference index, and updated GitHub Actions workflows to current (Node 20-compatible) action versions.