ncaa_bbStats is an open-source Python package for retrieving, parsing, and analyzing college baseball data: NCAA Division I, II, and III team statistics (2002–2026), player statistics (2021–2026), MLB Draft history (1965–2026), draft detail with signing bonuses (2021–2026), RPI and schedule strength, program finances, and a draft-prediction model with scouting reports.
Built for analysts, developers, and fans. Everything is cached locally, so it works offline; scraping is opt-in.
The draft model in this package powers a public site where you can browse a board, look up a player, or score a stat line of your own: https://codemateo15-ncaa-draft-app.share.connect.posit.cloud/
That site lives in a separate repository, which is a distinct project and is not covered by this package's MIT licence — it carries no licence, so all rights are reserved. This package is MIT; the app is not.
Note This project is under active development.
Documentation: ncaa_bbStats on ReadTheDocs
PyPI: ncaa-bbStats
Data sources and terms: DATA_PROVENANCE.md
pip install ncaa_bbStats # everything except predictions
pip install "ncaa_bbStats[model]" # + draft predictions
pip install "ncaa_bbStats[explain]" # + SHAP explanations
pip install "ncaa_bbStats[scrape]" # + re-scraping the sources yourselfRequires Python 3.10 or later.
from ncaa_bbStats import (
team_profile, leaderboard, scouting_report, resolve_team, luckiest_teams,
)
# Everything about one program in one season, across every dataset
p = team_profile("Tennessee", 2024)
p["record"] # 60-13, .822
p["rpi"]["rpi_rank"] # 1
p["draft"]["picks"] # 8
p["pythagorean"] # expected .807 against an actual .822
# Leaderboards that sort the right way round
leaderboard("era", stat_type="pitching", year=2025, min_ip=60, n=10)
leaderboard("cwrc+", year=2025, conference="SEC", n=10)
leaderboard("hr", per="career", qualifier="noMin", n=5)
# Every source spells schools differently; one id resolves them all
resolve_team("Eastern Ill.") == resolve_team("EIU") == resolve_team("Eastern Illinois")
# Who won more than their run differential deserved?
luckiest_teams(2025, n=5)
# A scouting report
print(scouting_report("Kade Anderson", 2025))| Dataset | Coverage |
|---|---|
| NCAA team statistics | 2002–2026, Divisions I–III |
| Player statistics | 2021–2026, Division I |
| MLB Draft history | 1965–2026, 69,781 picks |
| MLB Draft detail (bonuses, slots, biography) | 2021–2026, 3,685 picks |
| RPI, strength of schedule, quadrant records | 2021–2026, Division I |
| Program finances (EADA) | 2021–2025, carried forward to 2026 |
| Draft prospect rankings | 2021–2026 |
| Team registry | 1,023 programs |
| Player registry | 29,743 players over 61,279 player-seasons, 2021–2026 (no playing-time minimum) |
get_team_stat, display_team_stats, display_specific_team_stat,
list_all_teams, plot_team_stat_over_years, average_all_team_stats,
average_team_stat_str, average_team_stat_float
One canonical team_id per program, so datasets that spell schools differently
can be joined. Keyed on the federal IPEDS unitid where known, which survives
rebrands — Dixie State and Utah Tech share an id. Division is a per-season
attribute, not part of identity.
resolve_team, resolve_team_verbose, team_info, team_aliases,
team_seasons, team_division, team_conference, list_teams,
list_conferences, crosswalk
list_players, list_batters, list_pitchers, player_seasons,
batting_stat, pitching_stat, get_player_rows, load_player_frame,
list_available_years
The cache stores counting statistics only. Every rate and advanced statistic is computed when you read it, from those counts plus league constants this package derives from its own NCAA team data — so they can never fall out of step.
cwoba, cwraa, cwrc, cwrc_plus, cwsb, cspd, cfip, clob_pct,
league_constants, seasons_with_constants
College-calibrated analogues of the familiar sabermetric statistics, built the same way but with league constants regressed from NCAA play rather than borrowed from elsewhere. See DATA_PROVENANCE.md for the method and measured correlations.
leaderboard, stat_direction, qualification_rules
Takes the sort direction from the statistic, so a top-ERA list contains good pitchers. Supports playing-time floors, team and conference filters, and career aggregation that rebuilds rates from summed components.
parse_mlb_draft, get_drafted_players_mlb, get_drafted_players_college,
print_draft_picks_mlb, print_draft_picks_college (1965–2026)
draft_pick, draft_class, draft_history, slot_value, signing_bonus,
bonus_vs_slot, overslot_picks, biggest_bonuses, draft_demographics,
conference_draft_counts, state_pipeline (2021–2026, with bonuses and slots)
prospect_rank, prospect_board, prospect_vs_actual, biggest_draft_risers,
biggest_draft_fallers
rpi_rank, strength_of_schedule, rpi_table, rpi_record,
quadrant_record, home_road_neutral, nonconference_profile,
rpi_over_years, best_wins
program_finance, budget_percentile, roster_size, coaching_staff_size,
richest_programs, conference_spending, finance_vs_rpi
get_pythagorean_expectation, compare_pythagorean_expectation,
luck_rating, luckiest_teams, unluckiest_teams, pythagorean_exponent,
conference_exponents
team_profile, player_profile, draft_yield, dollars_per_draft_pick,
conference_report, pipeline, compare_teams
scouting_report, predict_draft_probability, predict_draft_order,
draft_board, explain_prediction, predict_from_stats, is_draft_eligible,
model_card
Two models: whether a player-season leads to being drafted (PR-AUC 0.602
against a 4.2% base rate -- a 14x lift over chance) and where a drafted player
falls in their class (Spearman 0.598 over 2,565 drafted players). Both are validated
leave-one-season-out across 2021-2026 -- each season is scored by a model
fitted without it -- and
draft_board serves those same out-of-fold predictions, so what the package
reports and what it shows you are the same numbers. Explanations come from SHAP
where installed, with a gain-based fallback that says which it used.
Read model_card() before quoting any of it — it carries the limitations,
including that Stage 1 precision depends on the base rate you apply it to (4.2%
here, over a population with no playing-time minimum, so these figures are not
comparable with a model scored over a qualified leaderboard), that eligibility
is inferred rather than looked up, and that order predictions separate tiers
rather than picks.
from ncaa_bbStats import predict_from_stats
result = predict_from_stats(
"pitcher", school_class="Jr",
stats={"era_pitch": 2.40, "so_pitch": 130, "bb_pitch": 25, "ip_pitch": 95.0},
team="LSU", season=2025,
)
print(result["report"])
result["confidence"] # 'low' -- reports how much had to be imputed- Team stat abbreviations
- Player stat abbreviations
- Team registry — how team names resolve
- Data provenance — sources, terms, and known limitations
Runnable notebooks covering every public function, with outputs saved so they
read without executing anything, live in notebooks/:
pip install -e ".[all]" jupyter
jupyter lab notebooks/Builders live in tools/ and the *_store modules; none of them ship in the
wheel. See tools/README.md.
python -m ncaa_bbStats.team_store --years 2026 # scrape NCAA team stats
python tools/build_league_constants.py # refit the run values
python tools/build_team_registry.py # rebuild the registry
python -m ncaa_bbStats.model_store # retrain the draft models
python -m pytest tests/ -q- Player statistics re-sourced from NCAA's own published data — partly done.
src/data/player_stats_cache_ncaa/covers 2021–2026 with no third-party export anywhere in the chain, and reproduces the default cache at r ≥ 0.998 on every column including this package's derived metrics. Read it withload_player_frame(..., source="ncaa"). It is not the default yet because NCAA publishes no date of birth (soageis empty), and its 2026 pitching rows carry now,l,cg,shoorsv— the upstream scrape did not collect them. Promoting it needs a retrain and a registry rebuild — see DATA_PROVENANCE.md - IPEDS identifiers backfilled for Division II and III programs
- Team game results with win-loss tracking
- Park factors, which currently limit
cwrc_plus
Found a bug or want a feature? Open an issue.
Star this repo and share to help support!
Mateo Biggs, mateojohn2024@gmail.com