Skip to content
Repoly

LLMKube

Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.

89 GOOD

Health breakdown: 89/100

Five terms, recomputed nightly from the GitHub API. How this is calculated.

  • Commit activity 35.0 / 35

    534 commits in 90 days (30+ scores full marks)

  • Release recency 25.0 / 25

    v0.9.27 released today

  • Stars 9.3 / 20

    210 GitHub stars (logarithmic)

  • Not archived 10.0 / 10

    Repository is active

  • Docker support 10.0 / 10

    Docker image or Compose file available

Commit activity

909 commits over the last 12 months
534
Last 90 days
909
Last 12 months
today
Last push

Ranked #12 of 14 tools in Generative Artificial Intelligence (GenAI) by health.

Details

LicenceApache-2.0
PlatformsGo, Docker, K8S
Latest releasev0.9.27 · today
Stars210
DockerYes
Repositorydefilantech/LLMKube ↗