STACKQUADRANT

omlx

jundot/omlx
8.2

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

Model Serving
20.8k1.8kPythonApache-2.0today

TensorRT-LLM

NVIDIA/TensorRT-LLM
7.2

TensorRT-LLM — a leading open-source project in the AI/LLM ecosystem.

Model Serving
14.5k2.7kPythonNOASSERTIONtoday

vllm-omni

vllm-project/vllm-omni
7.8

A framework for efficient model inference with omni-modality models

Model Serving
6.4k1.6kPythonApache-2.0today

Olares

beclab/Olares
7.4

Olares: An Open-Source Personal Cloud to Reclaim Your Data

Model Serving
5.2k321GoAGPL-3.0today

Deep-Learning-in-Production

ahkarami/Deep-Learning-in-Production
4.5

In this repository, I will share some useful notes and references about deploying deep learning-based models in production.

Model Serving
4.4k6851y ago

AI-Infra-from-Zero-to-Hero

HuaizhengZhang/AI-Infra-from-Zero-to-Hero
6.1

🚀 Awesome System for Machine Learning ⚡️ AI System Papers and Industry Practice. ⚡️ System for Machine Learning, LLM (Large Language Model), GenAI (Generative AI). 🍻 OSDI, NSDI, SIGCOMM, SoCC, MLSys, etc. 🗃️ Llama3, Mistral, etc. 🧑‍💻 Video Tutorials.

Model Serving
4.3k410MIT1y ago

LightLLM

ModelTC/LightLLM
7.1

LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.

Model Serving
4.2k356PythonApache-2.0today

ramalama

containers/ramalama
7.6

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

Model Serving
3.0k358PythonMIT3d ago

chitu

thu-pacman/chitu
7.1

High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

Model Serving
3.0k256PythonApache-2.0today

sie

superlinked/sie
6.9

Superlinked Inference Engine is an Open-source inference server and production cluster for embeddings, reranking, and extraction.

Model Serving
2.8k280PythonApache-2.0today

vllm-ascend

vllm-project/vllm-ascend
7.6

Community maintained hardware plugin for vLLM on Ascend

Model Serving
2.7k2.1kC++Apache-2.0today

inference

roboflow/inference
7.4

Turn any computer or edge device into a command center for your computer vision projects.

Model Serving
2.4k311PythonNOASSERTIONtoday

envd

tensorchord/envd
6.7

🏕️ Reproducible development environment for humans and agents

Model Serving
2.2k168GoApache-2.01mo ago

aici

microsoft/aici
4.8

AICI: Prompts as (Wasm) Programs

Model Serving
2.1k86RustMIT1y ago

mlrun

mlrun/mlrun
7.3

MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.

Model Serving
1.7k317PythonApache-2.01d ago

vllm-mlx

waybarrios/vllm-mlx
7.0

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

Model Serving
1.6k219PythonApache-2.01d ago

kitops

kitops-ml/kitops
7.0

An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.

Model Serving
1.4k185GoApache-2.03d ago

rtp-llm

alibaba/rtp-llm
6.1

RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

Model Serving
1.3k268CudaApache-2.0today

hopsworks

logicalclocks/hopsworks
5.8

Hopsworks - Data-Intensive AI platform with a Feature Store

Model Serving
1.3k160JavaAGPL-3.01y ago

truss

basetenlabs/truss
6.9

The simplest way to serve AI/ML models in production

Model Serving
1.2k122PythonMITtoday

TurboOCR

aiptimizer/TurboOCR
5.9

Fast GPU OCR server. 270 img/s on FUNSD. TensorRT FP16, PP-OCRv5, HTTP + gRPC.

Model Serving
1.0k100C++MIT7d ago

Nanoflow

efeslab/Nanoflow
4.6

A throughput-oriented high-performance serving framework for LLMs

Model Serving
97452Jupyter Notebook5mo ago

sglang-omni

sgl-project/sglang-omni
6.6

SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.

Model Serving
963394PythonApache-2.0today

model_server

openvinotoolkit/model_server
6.5

A scalable inference server for models optimized with OpenVINO™

Model Serving
920271C++Apache-2.0today

mosec

mosecorg/mosec
6.4

A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine

Model Serving
90273PythonApache-2.018d ago

pipeless

pipeless-ai/pipeless
5.0

An open-source computer vision framework to build and deploy apps in minutes

Model Serving
85252RustApache-2.02y ago

Yatai

bentoml/Yatai
5.9

Model Deployment at Scale on Kubernetes 🦄️

Model Serving
84076TypeScriptNOASSERTION2mo ago

ServerlessLLM

ServerlessLLM/ServerlessLLM
5.6

Serverless LLM Serving for Everyone.

Model Serving
71076PythonApache-2.03mo ago

timber

kossisoroyce/timber
5.2

Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference code. One command to load, one command to serve. 336x faster than Python inference.

Model Serving
68823PythonNOASSERTION4mo ago

fastapi-ml-skeleton

eightBEC/fastapi-ml-skeleton
4.4

FastAPI Skeleton App to serve machine learning models production-ready.

Model Serving
60492PythonApache-2.07mo ago

pinferencia

underneathall/pinferencia
4.7

Python + Inference - Model Deployment library in Python. Simplest model inference server ever.

Model Serving
54483PythonApache-2.03y ago

ome

ome-projects/ome
6.2

Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton

Model Serving
49592GoApache-2.0today

JetStream

AI-Hypercomputer/JetStream
4.8

JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).

Model Serving
45567PythonApache-2.07mo ago

xFasterTransformer

intel/xFasterTransformer
4.2

xFasterTransformer — open-source AI/LLM project.

Model Serving
43576C++Apache-2.011mo ago

gpu-rest-engine

NVIDIA/gpu-rest-engine
3.7

A REST API for Caffe using Docker and Go

Model Serving
42295C++BSD-3-Clause8y ago

stable-diffusion-deploy

Lightning-Universe/stable-diffusion-deploy
4.6

Learn to serve Stable Diffusion models on cloud infrastructure at scale. This Lightning App shows load-balancing, orchestrating, pre-provisioning, dynamic batching, GPU-inference, micro-services working together via the Lightning Apps framework.

Model Serving
39138PythonApache-2.02y ago

pmetal

Epistates/pmetal
4.8

PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.

Model Serving
31025RustNOASSERTION2mo ago

podman-desktop-extension-ai-lab

containers/podman-desktop-extension-ai-lab
5.8

Work with LLMs on a local environment using containers

Model Serving
29884TypeScriptApache-2.024d ago

BMW-YOLOv4-Inference-API-GPU

BMW-InnovationLab/BMW-YOLOv4-Inference-API-GPU
4.1

This is a repository for an nocode object detection inference API using the Yolov3 and Yolov4 Darknet framework.

Model Serving
27568PythonBSD-3-Clause4y ago

ggrun

raketenkater/ggrun
5.1

Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placement, measured flag tuning (AI Tune), hardware-matched HuggingFace downloads, and crash recovery. An Ollama alternative for multi-GPU rigs.

Model Serving
26715GoMITtoday

llm-server

raketenkater/llm-server
4.9

Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placement, measured flag tuning (AI Tune), hardware-matched HuggingFace downloads, and crash recovery. An Ollama alternative for multi-GPU rigs.

Model Serving
26715GoMITtoday

BMW-YOLOv4-Inference-API-CPU

BMW-InnovationLab/BMW-YOLOv4-Inference-API-CPU
3.9

This is a repository for an nocode object detection inference API using the Yolov4 and Yolov3 Opencv.

Model Serving
21759PythonNOASSERTION4y ago