STACKQUADRANT

onejune2018/Awesome-LLM-Eval

Evaluation & Testing

Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.

4.5
GitHub Metrics
Stars
658
Forks
84
Open Issues
48
Watchers
8
Contributors
5
Weekly Commits
0
Language
License
MIT
Last Commit
Nov 24, 2025
Created
Apr 26, 2023
Latest Release
Release Date
Synced: Aug 28, 2026
Quality Scores
Documentation Qualityw: 20%
5.7

Has docs site (https://arxiv.org/abs/2508.18646). Description: 198 chars. Stars signal: 658. Contributors: 5. Score: 5.7/10

Community Healthw: 20%
3.5

Stars: 658. Contributors: 5. Watchers: 8. Forks: 84. Issue ratio: 7.3%. Score: 3.5/10

Maintenance Velocityw: 15%
2.2

Last commit: 277d ago. Weekly commits: 0. No releases published. Maturity bonus: 3.3y old. Score: 2.2/10

API Design & DXw: 20%
6.4

Stars/issues ratio: 14. Has documentation site. Permissive license: MIT. Popularity signal: 658 stars. Score: 6.4/10

Production Readinessw: 15%
3.5

Battle-tested: 658 stars. Peer review: 5 contributors. No versioned releases. Licensed: MIT. Age: 3.3 years. Maintenance: last commit 277d ago. Score: 3.5/10

Ecosystem Integrationw: 10%
5.7

Fork interest: 84. Integration-friendly: MIT. Adoption: 658 stars. Has web presence. Score: 5.7/10

Tags
awsome-listawsome-listsbenchmarkbertchatglmchatgptdatasetevaluationgpt3large-language-model
Radar
Documentation Quality
Community Health
Maintenance Velocity
API Design & DX
Production Readiness
Ecosystem Integration