STACKQUADRANT

avifenesh/memra

Inference Engines

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

5.9
GitHub Metrics
Stars
327
Forks
38
Open Issues
2
Watchers
25
Contributors
7
Weekly Commits
250
Language
Rust
License
MIT
Last Commit
Aug 28, 2026
Created
Jul 5, 2026
Latest Release
v0.117.0
Release Date
Aug 28, 2026
Synced: Aug 28, 2026
Quality Scores
Documentation Qualityw: 20%
6.1

Has docs site (https://inference.tiyuvta.ai). Description: 265 chars. Stars signal: 325. Contributors: 7. Score: 6.1/10

Community Healthw: 20%
4.6

Stars: 325. Contributors: 7. Watchers: 25. Forks: 38. Issue ratio: 0.6%. Score: 4.6/10

Maintenance Velocityw: 15%
7.5

Last commit: 0d ago. Weekly commits: 0. Latest release: v0.116.1. Score: 7.5/10

API Design & DXw: 20%
7.9

Stars/issues ratio: 163. Typed language: Rust. Has documentation site. Permissive license: MIT. Popularity signal: 325 stars. Score: 7.9/10

Production Readinessw: 15%
3.5

Battle-tested: 325 stars. Peer review: 7 contributors. Versioned: v0.116.1. Licensed: MIT. Age: 0.1 years. Maintenance: last commit 0d ago. Score: 3.5/10

Ecosystem Integrationw: 10%
5.3

Fork interest: 38. Major ecosystem: Rust. Integration-friendly: MIT. Adoption: 325 stars. Has web presence. Score: 5.3/10

Tags
blackwellcudagemmaggufgpu-kernelsinference-enginellmllm-inferencellm-servingmoe
Radar
Documentation Quality
Community Health
Maintenance Velocity
API Design & DX
Production Readiness
Ecosystem Integration