Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
ci
ci-cd
cicd
evaluation
evaluation-framework
llm
llm-eval
llm-evaluation
llm-evaluation-framework
llmops
pentesting
prompt-engineering
prompt-testing
prompts
rag
red-teaming
testing
vulnerability-scanners
Updated 2026-10-11 21:24:49 +00:00
IUMBTEMS (I Use My Brain To Express My Self): High-integrity dialectic research agent harness for Claude Code, Pi (pi.dev), and OpenCode V2. Enforces verified empirical citations, dialectic thesis/antithesis swarms, and Socratic grilling.
ai-agents
autonomous-agents
citation-validation
claude-code
cli
deep-research
dialectic
epistemic-integrity
fact-checking
hallucination-prevention
mcp-servers
model-context-protocol
multi-agent-systems
opencode
opencode-plugin
pi-package
red-teaming
research-agent
socratic-questioning
swarm-intelligence
Updated 2026-10-11 21:07:24 +00:00
Adversary Emulation Framework
adversarial-attacks
adversary-simulation
c2
command-and-control
dns
dns-server
golang
gplv3
http
implant
red-team
red-team-engagement
red-teaming
security-tools
sliver
Updated 2026-10-10 17:05:26 +00:00
Complete Mandiant Offensive VM (Commando VM), a fully customizable Windows-based pentesting virtual machine distribution. [email protected]
Updated 2025-10-16 03:59:37 +00:00