A Python toolkit designed to scrape top GitHub user rankings from Gitstar Ranking, enrich user data with detailed repository metrics (owned vs. forked) via the GitHub REST API, and generate an automated leaderboard inside README.md.
All data files are automatically timestamped by year and month (YYYY_Mon), ensuring monthly historical tracking without data loss.
Last Updated:
2026_Sep(Extracted fromdata/gitstar_users_with_repo_counts_2026_Sep.csv)
| # | Username | Owned Repos (Sources) | GitHub_Stars | Gitstar Profile |
|---|---|---|---|---|
| 1 | vim-scripts | 5208 | 21667 | Profile |
| 2 | Apress | 3560 | 46191 | Profile |
| 3 | mattn | 1157 | 58938 | Profile |
| 4 | sindresorhus | 1130 | 1091868 | Profile |
| 5 | camenduru | 1074 | 38068 | Profile |
| 6 | nirzaf | 975 | 22879 | Profile |
| 7 | keijiro | 921 | 105582 | Profile |
| 8 | jonschlinkert | 797 | 27131 | Profile |
| 9 | egoist | 756 | 79544 | Profile |
| 10 | mafintosh | 669 | 50503 | Profile |
| 11 | simonw | 653 | 66037 | Profile |
| 12 | IonicaBizau | 538 | 25489 | Profile |
| 13 | WebReflection | 528 | 26683 | Profile |
| 14 | LaravelDaily | 501 | 30648 | Profile |
| 15 | schollz | 472 | 78021 | Profile |
| 16 | mattdesl | 447 | 39214 | Profile |
| 17 | swyxio | 447 | 26745 | Profile |
| 18 | thlorenz | 432 | 21107 | Profile |
| 19 | max-mapper | 413 | 44631 | Profile |
| 20 | kentcdodds | 403 | 54478 | Profile |
gitstar-ranking-scraper/
β
βββ assets/ # Project banners and visual assets
β βββ banner.svg
βββ data/ # Output directory for timestamped CSV datasets
β βββ gitstar_users_top10pages_2026_Sep.csv
β βββ gitstar_users_with_repo_counts_2026_Sep.csv
β
βββ scrape_gitstar_users.py # Step 1: Scrapes top 10 pages from Gitstar Ranking
βββ fetch_user_repo_counts.py # Step 2: Enriches dataset with GitHub repo breakdown
βββ update_readme_leaderboard.py # Step 3: Updates README leaderboard table
β
βββ .env # Environment variables (GitHub ADMIN_TOKEN)
βββ .gitignore # Specifies intentionally untracked files
βββ README.md # Project documentation & leaderboard
Scrapes user rankings directly from https://gitstar-ranking.com/users.
- Target Pages: First 10 pages (100 users per page = 1,000 top ranked users).
- Extracted Fields:
rank,username,stars,avatar_url,profile_url. - Features:
- Incremental Saving: Flushes progress to CSV immediately after each page is scraped.
- Resume Support: Skips already scraped usernames if interrupted.
- Timestamped Output: Saves to
data/gitstar_users_top10pages_YYYY_Mon.csv.
Enriches the scraped user list by fetching granular repository counts from the GitHub REST API.
- Metrics Collected:
sources_count: Owned / original repositories created by the user.forked_count: Repositories forked from other projects.total_repos_count: Total public repositories count.
- Features:
- Unauthenticated & Authenticated Fallback: Starts with unauthenticated API calls. If rate limit HTTP
403/429is reached, it seamlessly falls back toADMIN_TOKENloaded from.env. - Batch Saving: Saves progress to disk every 100 users.
- Resume Support: Reads existing enriched dataset to avoid redundant API requests upon rerun.
- Timestamped Output: Reads
data/gitstar_users_top10pages_YYYY_Mon.csvand outputsdata/gitstar_users_with_repo_counts_YYYY_Mon.csv.
- Unauthenticated & Authenticated Fallback: Starts with unauthenticated API calls. If rate limit HTTP
Automates updating the leaderboard table inside README.md.
- Features:
- Auto-Discovery: Automatically scans
data/and identifies the latest CSV dataset by date suffix (YYYY_Mon). - Sorting: Sorts users descending by
sources_count. - Leaderboard Rendering: Renders the Markdown table with links to profiles and updates
README.md.
- Auto-Discovery: Automatically scans
- Python 3.8+
requestsbeautifulsoup4
pip install requests beautifulsoup4To avoid GitHub API rate limits (60 requests/hour unauthenticated vs. 5,000 requests/hour authenticated), configure your GitHub Personal Access Token in .env:
ADMIN_TOKEN=github_pat_your_token_hereRun the pipeline sequentially using standard Python:
# Step 1: Scrape top 10 pages from Gitstar Ranking
python scrape_gitstar_users.py
# Step 2: Fetch repo counts (owned vs forked) via GitHub API
python fetch_user_repo_counts.py
# Step 3: Update README.md leaderboard table
python update_readme_leaderboard.pyBecause all output files are automatically timestamped with YYYY_Mon (e.g., 2026_Sep), running the pipeline each month creates a clean historical record inside data/ without overwriting prior months.
To change the number of pages scraped:
- Open
scrape_gitstar_users.py. - Update
num_pages:scrape_gitstar_users(num_pages=20) # Scrape top 20 pages (2,000 users)
To change how frequently progress is written during repo count fetching:
- Open
fetch_user_repo_counts.py. - Modify
batch_size:process_users(batch_size=50) # Flushes progress every 50 users
If fetch_user_repo_counts.py encounters rate limit warnings:
- Verify
ADMIN_TOKENin.envis valid and active. - Check token permissions (
public_repoor fine-grained read access).
Thank you for checking out this project! If you find this toolkit useful, please consider supporting its ongoing development:
- π Star this repository to show your support!
- π΄ Fork it to customize and add new features.
- π’ Share it with your fellow developers and community!
- β Buy me a coffee / Sponsor: Feel free to support via GitHub Sponsors.
Generated automatically by update_readme_leaderboard.py.