STACKQUADRANT

deepeval

confident-ai/deepeval
8.3

The LLM Evaluation Framework

Evaluation & Testing
17.9k1.9kPythonApache-2.01d ago

Ragas

explodinggradients/ragas
7.2

Ragas — a leading open-source project in the AI/LLM ecosystem.

Evaluation & Testing
15.5k1.7kPythonApache-2.06mo ago

iFixAi

ifixai-ai/iFixAi
7.6

The open-source diagnostic for AI misalignment. 32 tests across fabrication, manipulation, deception, unpredictability, and opacity. Provider-agnostic. Runs against OpenAI, Anthropic, Bedrock, Azure, Gemini, and more. Letter grade in under 5 minutes, content-addressed manifest for bit-identical replay. Built by iMe.

Evaluation & Testing
11.3k1.2kPythonApache-2.0today

garak

NVIDIA/garak
7.7

the LLM vulnerability scanner

Evaluation & Testing
9.0k1.2kPythonApache-2.02d ago

chinese-llm-benchmark

jeinlee1991/chinese-llm-benchmark
6.3

ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括335个大模型,覆盖chatgpt、gpt-5.2、o4-mini、谷歌gemini-3-pro、Claude-4.5、文心ERNIE-X1.1、ERNIE-5.0-Thinking、qwen3-max、百川、讯飞星火、商汤senseChat等商用模型, 以及kimi-k2、ernie4.5、minimax-M2、deepseek-v3.2、qwen3-2507、llama4、智谱GLM-4.6、gemma3、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。

Evaluation & Testing
6.4k2624d ago

AI-Infra-Guard

Tencent/AI-Infra-Guard
7.8

A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.

Evaluation & Testing
6.0k552PythonApache-2.0today

ouroboros

Q00/ouroboros
7.8

Agent OS: Stop prompting. Start specifying. A Socratic interview gates the spec on an ambiguity score, then one command drives execution, a 3-stage evaluation gate, and a budgeted evolution loop. MCP server, 13 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.

Evaluation & Testing
5.7k570PythonMITtoday

LLM-Engineers-Handbook

PacktPublishing/LLM-Engineers-Handbook
6.4

The LLM's practical guide: From the fundamentals to deploying advanced LLM and RAG apps to AWS using LLMOps best practices

Evaluation & Testing
5.3k1.3kPythonMIT4mo ago

agenta

Agenta-AI/agenta
7.6

The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

Evaluation & Testing
4.6k645TypeScriptNOASSERTIONtoday

lmms-eval

EvolvingLMMs-Lab/lmms-eval
7.7

One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

Evaluation & Testing
4.4k647PythonNOASSERTIONtoday

trulens

truera/trulens
7.5

Evaluation and Tracking for LLM Experiments and AI Agents

Evaluation & Testing
3.5k333PythonMITtoday

lmnr

lmnr-ai/lmnr
7.1

Laminar - open-source observability platform purpose-built for AI agents. YC S24.

Evaluation & Testing
3.2k227TypeScriptApache-2.0today

Observal

BlazeUp-AI/Observal
6.3

Observal is an AI agent registry with first in class observabilty and eval framework

Evaluation & Testing
2.3k471PythonNOASSERTION2d ago

future-agi

future-agi/future-agi
7.0

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

Evaluation & Testing
1.8k552PythonApache-2.0today

aisheets

huggingface/aisheets
6.0

Build, enrich, and transform datasets using AI models with no code

Evaluation & Testing
1.6k139TypeScriptApache-2.03mo ago

FuzzyAI

cyberark/FuzzyAI
5.3

A powerful tool for automated LLM fuzzing. It is designed to help developers and security researchers identify and mitigate potential jailbreaks in their LLM APIs.

Evaluation & Testing
1.6k215Jupyter NotebookApache-2.06mo ago

passmark

bug0inc/passmark
5.9

The open-source Playwright library for AI browser regression testing with intelligent caching, auto-healing, and multi-model verification.

Evaluation & Testing
1.3k184TypeScriptNOASSERTION28d ago

prompty

microsoft/prompty
7.0

Prompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty is an asset class and format for LLM prompts designed to enhance observability, understandability, and portability for developers.

Evaluation & Testing
1.3k126RustMITtoday

uqlm

cvs-health/uqlm
6.7

UQLM: Uncertainty Quantification for Language Models, is a Python package for UQ-based LLM hallucination detection

Evaluation & Testing
1.2k131PythonApache-2.01d ago

Tracely-ai

Jwuthri/Tracely-ai
6.1

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

Evaluation & Testing
1.2k104PythonMITtoday

judgeval

JudgmentLabs/judgeval
6.6

The open source post-building layer for agents. Our environment data and evals power agent post-training (RL, SFT) and monitoring.

Evaluation & Testing
1.1k96PythonApache-2.04d ago

FinSight-AI

juanjuandog/FinSight-AI
5.4

AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.

Evaluation & Testing
1.0k61JavaMIT4d ago

WHartTest

MGdaasLab/WHartTest
6.6

WHartTest 是一款AI驱动的测试自动化平台,实现从需求到可执行测试用例的自动化生成与管理,帮助测试团队提升效率与覆盖率。 (WHartTest is an AI-driven test automation platform that automates the generation and management of executable test cases from requirements, helping testing teams improve efficiency and coverage.)

Evaluation & Testing
1.0k159PythonMITtoday

scenario

langwatch/scenario
6.2

Agentic testing for agentic codebases

Evaluation & Testing
95578PythonMITtoday

aimock

CopilotKit/aimock
6.6

Mock everything your AI app talks to — LLM APIs, MCP, A2A, vector DBs, search. One package, one port, zero dependencies.

Evaluation & Testing
90264TypeScriptMITtoday

agent-qa

vostride/agent-qa
5.6

The self-improving Agentic QA harness with Memory. Write tests in natural language.
 Catch regressions before releases ship.

Evaluation & Testing
86615TypeScriptNOASSERTION24d ago

awesome-evals

benchflow-ai/awesome-evals
5.1

A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.

Evaluation & Testing
84789NOASSERTION7d ago

agent-skills-eval

darkrishabh/agent-skills-eval
5.2

A test runner for agentskills.io-style AI agent skills

Evaluation & Testing
70635TypeScriptMIT22d ago

Awesome-LLM-Eval

onejune2018/Awesome-LLM-Eval
4.5

Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.

Evaluation & Testing
65884MIT9mo ago

Awesome-LLM-in-Social-Science

ValueByte-AI/Awesome-LLM-in-Social-Science
5.3

Awesome papers involving LLMs in Social Science.

Evaluation & Testing
64652MIT8d ago

langtest

PacificAI/langtest
6.3

Deliver safe & effective language models

Evaluation & Testing
55952PythonApache-2.07d ago

langtest

Pacific-AI-Corp/langtest
6.3

Deliver safe & effective language models

Evaluation & Testing
55952PythonApache-2.07d ago

fakecloud

faiscadev/fakecloud
6.0

Free, open-source AWS emulator. LocalStack alternative: 26 services, 1,924 operations, 100% conformance. No account, no auth token, no paid tier.

Evaluation & Testing
53540RustAGPL-3.03d ago

continuous-eval

relari-ai/continuous-eval
5.6

Data-Driven Evaluation for LLM-Powered Applications

Evaluation & Testing
51638PythonApache-2.017d ago

rhesis

rhesis-ai/rhesis
5.9

The testing platform for AI teams. Bring engineers, PMs, and domain experts together to generate tests, simulate (adversarial) conversations, and trace every failure to its root cause.

Evaluation & Testing
39032PythonNOASSERTIONtoday

flutter-skill

ai-dashboad/flutter-skill
5.5

AI-powered E2E testing for 10 platforms. 253 MCP tools. Zero config. Works with Claude, Cursor, Windsurf, Copilot. Test Flutter, React Native, iOS, Android, Web, Electron, Tauri, KMP, .NET MAUI — all from natural language.

Evaluation & Testing
36052DartMIT6d ago

llm-leaderboard

JonathanChavezTamales/llm-leaderboard
4.6

A comprehensive set of LLM benchmark scores and provider prices. (deprecated, read more in README)

Evaluation & Testing
35740JavaScriptNOASSERTION10mo ago

palico-ai

palico-ai/palico-ai
4.5

Build, Improve Performance, and Productionize your LLM Application with an Integrated Framework

Evaluation & Testing
34331TypeScriptMIT1y ago

llms-tools

PetroIvaniuk/llms-tools
4.8

A list of LLMs Tools & Projects

Evaluation & Testing
32750Apache-2.028d ago

athina-evals

athina-ai/athina-evals
4.0

Python SDK for running evaluations on LLM generated responses

Evaluation & Testing
30123Python1y ago

testdriverai

testdriverai/testdriverai
4.7

Computer-Use SDK for E2E QA Testing

Evaluation & Testing
24036JavaScript11d ago

qaskills

PramodDutta/qaskills
4.6

QA Skills Directory QA Skills is a curated directory of testing-specific skills for AI coding agents (Claude Code, Cursor, Copilot, etc.).

Evaluation & Testing
21022TypeScripttoday

awesome-qa-skills

naodeng/awesome-qa-skills
5.1

Awesome QA Skills — a bilingual (zh/en) AI testing Agent Skills library for Codex, Cursor, Claude Code, Kiro, OpenCode, and Trae. Ships 4 testing workflows and 25 testing-type skills (58 skill folders with language parity): independently installable, composable, and eval-ready with skill-up. Covers requirements, strategy, cases, API/performance/sec

Evaluation & Testing
18425PythonGPL-3.04d ago

ai-api-test-skill

buer2233/ai-api-test-skill
4.3

AI接口自动化测试 Skill:面向 Python + pytest + requests,驱动 Codex / Claude Code 生成、维护和调试接口用例(AI API test automation skill for Python + pytest + requests)

Evaluation & Testing
13815PythonMIT20d ago

spec2case

blackhaiyu-sudo/spec2case
4.1

Spec2Case 是生产级 AI 测试用例生成智能体,支持图片/文本需求理解、人工确认、LangGraph 流程编排和 Excel 用例导出。

Evaluation & Testing
1186PythonMIT3mo ago

xmind-vault

Circleoipillar/xmind-vault
2.1

XMind Vault

Evaluation & Testing
3410d ago