Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
ci
ci-cd
cicd
evaluation
evaluation-framework
llm
llm-eval
llm-evaluation
llm-evaluation-framework
llmops
pentesting
prompt-engineering
prompt-testing
prompts
rag
red-teaming
testing
vulnerability-scanners
Updated 2026-10-11 21:24:49 +00:00
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
agentops
agents
ai
ai-governance
apache-spark
evaluation
langchain
llm-evaluation
llmops
machine-learning
ml
mlflow
mlops
model-management
observability
open-source
openai
prompt-engineering
Updated 2026-10-11 21:10:33 +00:00
AI Observability & Evaluation
agents
ai-monitoring
ai-observability
aiengineering
anthropic
datasets
evals
langchain
llamaindex
llm-eval
llm-evaluation
llmops
llms
openai
prompt-engineering
smolagents
Updated 2026-10-11 21:10:33 +00:00
🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
analytics
autogen
evaluation
langchain
large-language-models
llama-index
llm
llm-evaluation
llm-observability
llmops
monitoring
observability
open-source
openai
playground
prompt-engineering
prompt-management
self-hosted
ycombinator
Updated 2026-10-11 02:54:33 +00:00