Evaluation and Tracking for LLM Experiments and AI Agents
agent-evaluation
agentops
ai-agents
ai-monitoring
ai-observability
evals
explainable-ml
llm-eval
llm-evaluation
llmops
llms
machine-learning
neural-networks
Updated 2026-10-11 21:24:50 +00:00
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
ci
ci-cd
cicd
evaluation
evaluation-framework
llm
llm-eval
llm-evaluation
llm-evaluation-framework
llmops
pentesting
prompt-engineering
prompt-testing
prompts
rag
red-teaming
testing
vulnerability-scanners
Updated 2026-10-11 21:24:49 +00:00
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
ai-gateway
anthropic
azure-openai
bedrock
gateway
langchain
litellm
llm
llm-gateway
llmops
mcp-gateway
openai
openai-proxy
rust
rust-ai
vertex-ai
Updated 2026-10-11 21:11:56 +00:00
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
agentops
agents
ai
ai-governance
apache-spark
evaluation
langchain
llm-evaluation
llmops
machine-learning
ml
mlflow
mlops
model-management
observability
open-source
openai
prompt-engineering
Updated 2026-10-11 21:10:33 +00:00
AI Observability & Evaluation
agents
ai-monitoring
ai-observability
aiengineering
anthropic
datasets
evals
langchain
llamaindex
llm-eval
llm-evaluation
llmops
llms
openai
prompt-engineering
smolagents
Updated 2026-10-11 21:10:33 +00:00
🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
analytics
autogen
evaluation
langchain
large-language-models
llama-index
llm
llm-evaluation
llm-observability
llmops
monitoring
observability
open-source
openai
playground
prompt-engineering
prompt-management
self-hosted
ycombinator
Updated 2026-10-11 02:54:33 +00:00
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
ai-inference
deep-learning
generative-ai
inference-platform
llm
llm-inference
llm-serving
llmops
machine-learning
ml-engineering
mlops
model-inference-service
model-serving
multimodal
python
Updated 2026-10-05 17:17:20 +00:00
Supercharge Your LLM Application Evaluations 🚀
Updated 2026-02-24 07:47:18 +00:00