A command-line benchmarking tool
Updated 2026-10-07 21:06:41 +00:00
Production-grade evaluation framework & turnkey GitHub Action for benchmarking AI coding harnesses (Claude Code, Gemini CLI, Antigravity, OpenCode, DeepSeek), Model Context Protocol (MCP) servers, LSP diagnostics, and agent plugins across CoderEval & Terminal-Bench.
Updated 2026-08-22 20:33:17 +00:00