coder_eval

UiPath/coder_eval
★ 107 stars Python AI/LLM Updated today
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
View on GitHub → 🔍 Audit Wallet Slippage →

Quick Install

Copy the config for your editor. Some servers may need additional setup — check the README.

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "coder_eval": {
      "command": "uvx",
      "args": [
        "coder-eval"
      ]
    }
  }
}

Or install with pip: pip install coder-eval

README Excerpt

**Coder Eval** (`pip install coder-eval` / `uv tool install coder-eval`) is an open-source framework for **evaluating and benchmarking AI coding agents and their skills** — built for CLI and skill builders — with sandboxing, reproducibility, and data-driven analysis. It runs a real agent (**Claude Code**, **Codex**, or **Google Antigravity /

Tools (5)

envmodeltagstasksversion

Topics

agent-evaluationagent-skillsagent-testinganthropicanthropic-claudeantigravityclaudeclaude-agent-sdkclaude-codeclaude-code-skillsclaude-skillscodexcoding-agentsevaluation-frameworkgemini