Model Runtime
4 tools.
local-runinference-servermodel-packaging
B
BentoML
inference-server · model-packagingAn open-source framework and platform designed for packaging, deploying, and scaling AI/ML models and custom inference pipelines in production.
inferencemodel-servingllm-inference
D
Dynamo-Triton
inference-serverAn open-source inference server designed to deploy AI models across major frameworks (TensorRT, PyTorch, ONNX, OpenVINO, Python, and RAPIDS FIL) on GPUs, CPUs, and accelerators.
model-servinginference-servermlops
O
Ollama
inference-server · local-runA tool that enables running open large language models locally and scaling to the cloud.
local-llmmodel-inferencecloud-models
V
vLLM
free · inference-serverA high-throughput and memory-efficient LLM inference and serving engine designed for easy, fast, and cost-efficient model deployment on multiple hardware platforms.
llm-servinginference-enginepagedattention