STACKQUADRANT

llama.cpp

ggml-org/llama.cpp
8.3

llama.cpp — a leading open-source project in the AI/LLM ecosystem.

Inference Engines
126.0k22.3kC++MITtoday

vLLM

vllm-project/vllm
8.6

vLLM — a leading open-source project in the AI/LLM ecosystem.

Inference Engines
90.2k21.3kPythonApache-2.0today

gpt4all

nomic-ai/gpt4all
7.1

GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.

Inference Engines
77.4k8.3kC++MIT1y ago

ray

ray-project/ray
8.6

Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

Inference Engines
43.6k8.0kPythonApache-2.0today

gitleaks

gitleaks/gitleaks
8.1

Find secrets with Gitleaks 🔑

Inference Engines
29.0k2.2kGoMIT1d ago

llm-action

liguodongiot/llm-action
6.7

本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)

Inference Engines
25.0k2.8kHTMLApache-2.01mo ago

litgpt

Lightning-AI/litgpt
7.9

20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.

Inference Engines
13.6k1.5kPythonApache-2.010d ago

Halfrost-Field

halfrost/Halfrost-Field
6.9

✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记

Inference Engines
13.2k1.9kGoCC-BY-SA-4.021d ago

OpenLLM

bentoml/OpenLLM
7.4

Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.

Inference Engines
12.5k835PythonApache-2.03d ago

mistral-inference

mistralai/mistral-inference
7.0

Official inference library for Mistral models

Inference Engines
10.8k1.1kJupyter NotebookApache-2.02mo ago

openvino

openvinotoolkit/openvino
8.2

OpenVINO™ is an open source toolkit for optimizing and deploying AI inference

Inference Engines
10.7k3.3kC++Apache-2.0today

PowerInfer

Tiiny-AI/PowerInfer
6.7

High-speed Large Language Model Serving for Local Deployment

Inference Engines
9.8k596C++MIT3mo ago

BentoML

bentoml/BentoML
8.0

The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

Inference Engines
8.8k1.0kPythonApache-2.01d ago

lmdeploy

InternLM/lmdeploy
7.8

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

Inference Engines
8.0k731PythonApache-2.0today

dynamo

ai-dynamo/dynamo
7.5

A Datacenter Scale Distributed Inference Serving Framework

Inference Engines
7.9k1.5kRustNOASSERTIONtoday

openevolve

algorithmicsuperintelligence/openevolve
6.9

Open-source implementation of AlphaEvolve

Inference Engines
7.3k1.1kPythonApache-2.01mo ago

plano

katanemo/plano
7.5

Plano is an AI-native proxy server and data plane for agentic apps - centralizing orchestration, safety, observability, and smart LLM routing so you can deliver agents faster.

Inference Engines
7.0k488RustApache-2.08d ago

kimi-k3-in-c

FareedKhan-dev/kimi-k3-in-c
7.3

A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

Inference Engines
6.6k1.1kCApache-2.01d ago

turbo-fieldfare

drumih/turbo-fieldfare
6.4

Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

Inference Engines
6.4k402SwiftApache-2.0today

flashinfer

flashinfer-ai/flashinfer
7.9

FlashInfer: Kernel Library for LLM Serving

Inference Engines
6.3k1.3kPythonApache-2.0today

kserve

kserve/kserve
8.0

Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes

Inference Engines
5.8k1.6kGoApache-2.0today

shimmy

Michael-A-Kuykendall/shimmy
6.3

⚡ Python-free Rust inference server — OpenAI-API compatible. GGUF + SafeTensors, hot model swap, auto-discovery, single binary. FREE now, FREE forever.

Inference Engines
5.8k561RustApache-2.0today

gpustack

gpustack/gpustack
7.4

Performance-optimized AI inference on your GPUs. Unlock superior throughput by selecting and tuning engines like vLLM or SGLang.

Inference Engines
5.6k628PythonApache-2.0today

lemonade

lemonade-sdk/lemonade
7.6

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

Inference Engines
5.5k479C++Apache-2.0today

Awesome-LLM-Inference

xlite-dev/Awesome-LLM-Inference
6.6

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

Inference Engines
5.5k429PythonGPL-3.013d ago

eko

FellouAI/eko
6.7

Eko (Eko Keeps Operating) - Build Production-ready Agentic Workflow with Natural Language - eko.fellou.ai

Inference Engines
5.0k442TypeScriptMIT5mo ago

ruvector

ruvnet/ruvector
7.0

RuVector is a High Performance, Real-Time, Self-Learning, Vector Graph Neural Network, and Database built in Rust.

Inference Engines
4.5k590RustMITtoday

RuVector

ruvnet/RuVector
7.3

RuVector is a High Performance, Real-Time, Self-Learning, Vector Graph Neural Network, and Database built in Rust.

Inference Engines
4.5k590RustMITtoday

optillm

algorithmicsuperintelligence/optillm
6.6

Optimizing inference proxy for LLMs

Inference Engines
4.3k384PythonApache-2.01mo ago

lorax

predibase/lorax
6.6

Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

Inference Engines
3.8k326PythonApache-2.03mo ago

AI-Engineer-Headquarters

hemansnation/AI-Engineer-Headquarters
5.3

A collection of scientific methods, processes, algorithms, and systems to build stories & models.

Inference Engines
3.7k699Jupyter Notebook9mo ago

deepsparse

neuralmagic/deepsparse
5.9

Sparsity-aware deep learning inference runtime for CPUs

Inference Engines
3.2k193PythonNOASSERTION1y ago

spiceai

spiceai/spiceai
7.3

A portable accelerated SQL query, search, and LLM-inference engine, written in Rust, for data-grounded AI apps and agents.

Inference Engines
3.1k223RustApache-2.0today

distributed-llama

b4rtaz/distributed-llama
6.1

Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.

Inference Engines
3.0k247C++MIT1mo ago

Medusa

FasterDecoding/Medusa
5.5

Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads

Inference Engines
2.8k205Jupyter NotebookApache-2.02y ago

kvcached

ovg-project/kvcached
5.9

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

Inference Engines
1.1k133PythonApache-2.04d ago

nobodywho

nobodywho-ooo/nobodywho
6.8

NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.

Inference Engines
1.1k77RustEUPL-1.2today

mlxstudio

jjang-ai/mlxstudio
5.4

MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)

Inference Engines
95965today

ZhiLight

zhihu/ZhiLight
5.1

A highly optimized LLM inference acceleration engine for Llama and its variants.

Inference Engines
908104C++Apache-2.05mo ago

pegainfer

pegainfer-project/pegainfer
6.5

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

Inference Engines
657103RustApache-2.01d ago

openinfer

openinfer-project/openinfer
6.5

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

Inference Engines
657103RustApache-2.01d ago

yalm

andrewkchan/yalm
3.6

Yet Another Language Model: LLM inference in C++/CUDA, no libraries except for I/O

Inference Engines
59664C++11mo ago

KuiperLLama

zjhellofss/KuiperLLama
3.9

校招、秋招、春招、实习好项目,带你从零动手实现支持LLama2/3和Qwen2.5的大模型推理框架。

Inference Engines
571143C++10mo ago

tessera

zengxiao-he/tessera
4.3

From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.

Inference Engines
5239PythonNOASSERTION2mo ago

TensorSharp

zhongkaifu/TensorSharp
6.0

A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability

Inference Engines
37738C#BSD-3-Clausetoday

swiftLLM

interestingLSY/swiftLLM
3.8

A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).

Inference Engines
33337PythonApache-2.01y ago

memra

avifenesh/memra
5.9

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

Inference Engines
32638RustMITtoday