All categories

Model Runtime

4 tools.

local-runinference-servermodel-packaging
B

BentoML

inference-server · model-packaging

An open-source framework and platform designed for packaging, deploying, and scaling AI/ML models and custom inference pipelines in production.

inferencemodel-servingllm-inference
D

Dynamo-Triton

inference-server

An open-source inference server designed to deploy AI models across major frameworks (TensorRT, PyTorch, ONNX, OpenVINO, Python, and RAPIDS FIL) on GPUs, CPUs, and accelerators.

model-servinginference-servermlops
O

Ollama

inference-server · local-run

A tool that enables running open large language models locally and scaling to the cloud.

local-llmmodel-inferencecloud-models
V

vLLM

free · inference-server

A high-throughput and memory-efficient LLM inference and serving engine designed for easy, fast, and cost-efficient model deployment on multiple hardware platforms.

llm-servinginference-enginepagedattention