Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
ci
ci-cd
cicd
evaluation
evaluation-framework
llm
llm-eval
llm-evaluation
llm-evaluation-framework
llmops
pentesting
prompt-engineering
prompt-testing
prompts
rag
red-teaming
testing
vulnerability-scanners
Updated 2026-10-11 21:24:49 +00:00
The LLM Evaluation Framework
evaluation-framework
evaluation-metrics
llm-evaluation
llm-evaluation-framework
llm-evaluation-metrics
python
Updated 2026-10-10 00:15:09 +00:00