LLMOps
8 tools.
Braintrust
eval · observabilityAn AI observability and evaluation platform designed to monitor, trace, and test AI products and agents at scale.
Confident AI
eval · freeAn enterprise AI evaluation, red teaming, and observability platform that helps teams monitor, trace, and stress-test LLM applications.
Galileo AI
eval · guardrailsAn enterprise-grade AI observability and evaluation platform for LLM applications, enabling developers to run offline evaluations, auto-tune metrics, and deploy optimized low-latency guardrails using compact Luna models.
Guardrails AI
eval · guardrailsAn AI reliability platform and framework designed for building, governing, and scaling production GenAI applications. It provides runtime guardrails to detect policy violations, hallucinations, and data leakage, alongside capabilities for simulating realistic datasets and generating evaluation datasets.
Helicone
observabilityAn AI gateway and LLM observability platform designed to route, debug, and analyze LLM applications.
Langfuse
eval · observabilityAn open-source AI engineering platform designed for LLM observability, tracing, prompt management, evaluation, and experiments.
Ragas
eval · freeAn open-source framework designed to evaluate and monitor Retrieval-Augmented Generation (RAG) systems by providing automated metrics, synthetic test data generation, and online quality tracking.
TruLens
eval · freeAn open-source evaluation and tracing framework for AI agents and LLM applications, allowing developers to measure execution flows and application quality using metrics like groundedness and context relevance.