evaluation

This document provides guidance on creating comprehensive evaluations for MCP servers.

MCP Server Evaluation Guide

Overview

This document provides guidance on creating comprehensive evaluations for MCP servers. Evaluations test whether LLMs can effectively use your MCP server to answer realistic, complex questions using only the tools provided.


Quick Reference

Evaluation Requirements

Output Format

<evaluation>
   <qa_pair>
      <question>Your question here</question>
      <answer>Single verifiable answer</answer>
   </qa_pair>
</evaluation>

Purpose of Evaluations

The measure of quality of an MCP server is NOT how well or comprehensively the server implements tools, but how well these implementations (input/output schemas, docstrings/descriptions, functionality) enable LLMs with no other context and access ONLY to the MCP servers to answer realistic and difficult questions.

Evaluation Overview

Create 10 human-readable questions requiring ONLY READ-ONLY, INDEPENDENT, NON-DESTRUCTIVE, and IDEMPOTENT operations to answer. Each question should be:

Question Guidelines

Core Requirements

  1. Questions MUST be independent
  1. Questions MUST require ONLY NON-DESTRUCTIVE AND IDEMPOTENT tool use
  1. Questions must be REALISTIC, CLEAR, CONCISE, and COMPLEX

Complexity and Depth

  1. Questions must require deep exploration
  1. Questions may require extensive paging
  1. Questions must require deep understanding
  1. Questions must not be solvable with straightforward keyword search

Tool Testing

  1. Questions should stress-test tool return values
  1. Questions should MOSTLY reflect real human use cases
  1. Questions may require dozens of tool calls
  1. Include ambiguous questions

Stability

  1. Questions must be designed so the answer DOES NOT CHANGE
  1. DO NOT let the MCP server RESTRICT the kinds of questions you create

Answer Guidelines

Verification

  1. Answers must be VERIFIABLE via direct string comparison

Readability

  1. Answers should generally prefer HUMAN-READABLE formats

Stability

  1. Answers must be STABLE/STATIONARY
  1. Answers must be CLEAR and UNAMBIGUOUS

Diversity

  1. Answers must be DIVERSE
  1. Answers must NOT be complex structures

Evaluation Process

Step 1: Documentation Inspection

Read the documentation of the target API to understand:

Step 2: Tool Inspection

List the tools available in the MCP server:

Step 3: Developing Understanding

Repeat steps 1 & 2 until you have a good understanding:

Step 4: Read-Only Content Inspection

After understanding the API and tools, USE the MCP server tools:

Step 5: Task Generation

After inspecting the content, create 10 human-readable questions:

Output Format

Each QA pair consists of a question and an answer. The output should be an XML file with this structure:

<evaluation>
   <qa_pair>
      <question>Find the project created in Q2 2024 with the highest number of completed tasks. What is the project name?</question>
      <answer>Website Redesign</answer>
   </qa_pair>
   <qa_pair>
      <question>Search for issues labeled as "bug" that were closed in March 2024. Which user closed the most issues? Provide their username.</question>
      <answer>sarah_dev</answer>
   </qa_pair>
   <qa_pair>
      <question>Look for pull requests that modified files in the /api directory and were merged between January 1 and January 31, 2024. How many different contributors worked on these PRs?</question>
      <answer>7</answer>
   </qa_pair>
   <qa_pair>
      <question>Find the repository with the most stars that was created before 2023. What is the repository name?</question>
      <answer>data-pipeline</answer>
   </qa_pair>
</evaluation>

Evaluation Examples

Good Questions

Example 1: Multi-hop question requiring deep exploration (GitHub MCP)

<qa_pair>
   <question>Find the repository that was archived in Q3 2023 and had previously been the most forked project in the organization. What was the primary programming language used in that repository?</question>
   <answer>Python</answer>
</qa_pair>

This question is good because:

Example 2: Requires understanding context without keyword matching (Project Management MCP)

<qa_pair>
   <question>Locate the initiative focused on improving customer onboarding that was completed in late 2023. The project lead created a retrospective document after completion. What was the lead's role title at that time?</question>
   <answer>Product Manager</answer>
</qa_pair>

This question is good because:

Example 3: Complex aggregation requiring multiple steps (Issue Tracker MCP)

<qa_pair>
   <question>Among all bugs reported in January 2024 that were marked as critical priority, which assignee resolved the highest percentage of their assigned bugs within 48 hours? Provide the assignee's username.</question>
   <answer>alex_eng</answer>
</qa_pair>

This question is good because:

Example 4: Requires synthesis across multiple data types (CRM MCP)

<qa_pair>
   <question>Find the account that upgraded from the Starter to Enterprise plan in Q4 2023 and had the highest annual contract value. What industry does this account operate in?</question>
   <answer>Healthcare</answer>
</qa_pair>

This question is good because:

Poor Questions

Example 1: Answer changes over time

<qa_pair>
   <question>How many open issues are currently assigned to the engineering team?</question>
   <answer>47</answer>
</qa_pair>

This question is poor because:

Example 2: Too easy with keyword search

<qa_pair>
   <question>Find the pull request with title "Add authentication feature" and tell me who created it.</question>
   <answer>developer123</answer>
</qa_pair>

This question is poor because:

Example 3: Ambiguous answer format

<qa_pair>
   <question>List all the repositories that have Python as their primary language.</question>
   <answer>repo1, repo2, repo3, data-pipeline, ml-tools</answer>
</qa_pair>

This question is poor because:

Verification Process

After creating evaluations:

  1. Examine the XML file to understand the schema
  2. Load each task instruction and in parallel using the MCP server and tools, identify the correct answer by attempting to solve the task YOURSELF
  3. Flag any operations that require WRITE or DESTRUCTIVE operations
  4. Accumulate all CORRECT answers and replace any incorrect answers in the document
  5. Remove any <qa_pair> that require WRITE or DESTRUCTIVE operations

Remember to parallelize solving tasks to avoid running out of context, then accumulate all answers and make changes to the file at the end.

Tips for Creating Quality Evaluations

  1. Think Hard and Plan Ahead before generating tasks
  2. Parallelize Where Opportunity Arises to speed up the process and manage context
  3. Focus on Realistic Use Cases that humans would actually want to accomplish
  4. Create Challenging Questions that test the limits of the MCP server's capabilities
  5. Ensure Stability by using historical data and closed concepts
  6. Verify Answers by solving the questions yourself using the MCP server tools
  7. Iterate and Refine based on what you learn during the process

Running Evaluations

After creating your evaluation file, you can use the provided evaluation harness to test your MCP server.

Setup

  1. Install Dependencies

``bash pip install -r scripts/requirements.txt ``

Or install manually: ``bash pip install anthropic mcp ``

  1. Set API Key

``bash export ANTHROPIC_API_KEY=your_api_key_here ``

Evaluation File Format

Evaluation files use XML format with <qa_pair> elements:

<evaluation>
   <qa_pair>
      <question>Find the project created in Q2 2024 with the highest number of completed tasks. What is the project name?</question>
      <answer>Website Redesign</answer>
   </qa_pair>
   <qa_pair>
      <question>Search for issues labeled as "bug" that were closed in March 2024. Which user closed the most issues? Provide their username.</question>
      <answer>sarah_dev</answer>
   </qa_pair>
</evaluation>

Running Evaluations

The evaluation script (scripts/evaluation.py) supports three transport types:

Important:

1. Local STDIO Server

For locally-run MCP servers (script launches the server automatically):

python scripts/evaluation.py \
  -t stdio \
  -c python \
  -a my_mcp_server.py \
  evaluation.xml

With environment variables:

python scripts/evaluation.py \
  -t stdio \
  -c python \
  -a my_mcp_server.py \
  -e API_KEY=abc123 \
  -e DEBUG=true \
  evaluation.xml

2. Server-Sent Events (SSE)

For SSE-based MCP servers (you must start the server first):

python scripts/evaluation.py \
  -t sse \
  -u https://example.com/mcp \
  -H "Authorization: Bearer token123" \
  -H "X-Custom-Header: value" \
  evaluation.xml

3. HTTP (Streamable HTTP)

For HTTP-based MCP servers (you must start the server first):

python scripts/evaluation.py \
  -t http \
  -u https://example.com/mcp \
  -H "Authorization: Bearer token123" \
  evaluation.xml

Command-Line Options

usage: evaluation.py [-h] [-t {stdio,sse,http}] [-m MODEL] [-c COMMAND]
                     [-a ARGS [ARGS ...]] [-e ENV [ENV ...]] [-u URL]
                     [-H HEADERS [HEADERS ...]] [-o OUTPUT]
                     eval_file

positional arguments:
  eval_file             Path to evaluation XML file

optional arguments:
  -h, --help            Show help message
  -t, --transport       Transport type: stdio, sse, or http (default: stdio)
  -m, --model           Claude model to use (default: claude-3-7-sonnet-20250219)
  -o, --output          Output file for report (default: print to stdout)

stdio options:
  -c, --command         Command to run MCP server (e.g., python, node)
  -a, --args            Arguments for the command (e.g., server.py)
  -e, --env             Environment variables in KEY=VALUE format

sse/http options:
  -u, --url             MCP server URL
  -H, --header          HTTP headers in 'Key: Value' format

Output

The evaluation script generates a detailed report including:

Save Report to File

python scripts/evaluation.py \
  -t stdio \
  -c python \
  -a my_server.py \
  -o evaluation_report.md \
  evaluation.xml

Complete Example Workflow

Here's a complete example of creating and running an evaluation:

  1. Create your evaluation file (my_evaluation.xml):
<evaluation>
   <qa_pair>
      <question>Find the user who created the most issues in January 2024. What is their username?</question>
      <answer>alice_developer</answer>
   </qa_pair>
   <qa_pair>
      <question>Among all pull requests merged in Q1 2024, which repository had the highest number? Provide the repository name.</question>
      <answer>backend-api</answer>
   </qa_pair>
   <qa_pair>
      <question>Find the project that was completed in December 2023 and had the longest duration from start to finish. How many days did it take?</question>
      <answer>127</answer>
   </qa_pair>
</evaluation>
  1. Install dependencies:
pip install -r scripts/requirements.txt
export ANTHROPIC_API_KEY=your_api_key
  1. Run evaluation:
python scripts/evaluation.py \
  -t stdio \
  -c python \
  -a github_mcp_server.py \
  -e GITHUB_TOKEN=ghp_xxx \
  -o github_eval_report.md \
  my_evaluation.xml
  1. Review the report in github_eval_report.md to:

Troubleshooting

Connection Errors

If you get connection errors:

Low Accuracy

If many evaluations fail:

Timeout Issues

If tasks are timing out: