← Back to Blog

2026-09-225 min read • Measure

How to Test an MCP Server for Accuracy

How to score an MCP server with a written list of questions, the page each one must cite, and one question the sources should refuse.

An impression is not a test. "It seemed fine" does not tell you which question failed, and it does not tell the next person what to rerun after a docs edit.

Write the questions down. Next to each one, write the URL that should be cited and the fact the passage has to contain. Run them after indexing finishes. Score a hit only when both match.

What the sheet is

Five to ten questions from real tasks. Three will show a broken server. Ten will show which task fails. Write them the way a person asks, not the way the heading is phrased. Include one question the sources do not answer, and score a hit when the server says so.

That sheet is yours. It is not a score you can paste onto someone else's server. A fuller way to mark an answer, with a worked comparison, is the MCP server rubric. Use the sheet every time you change a page. Use the rubric when pass or fail is too coarse.

Why both the link and the fact have to match

A fluent paragraph with the wrong link fails. The model can restate a neighboring page in clean sentences.

A nearby paragraph fails the out-of-scope question. The wifi password is a pass only when the server says the content does not cover it.

A reworded question is not a rerun. Changing the wording until a single row passes teaches you a prompt.

Run the sheet against your MCP Studio server

MCP Studio is the no-code builder on Appa Tools. The server has to exist and show indexing complete before the sheet means anything. Open it on the dashboard. If you still need the server, create it in the wizard.

Question Expected Pass if
Which header authenticates an upload? Authentication page The citation is that page and the header name matches
What is the maximum upload size? Limits page The number matches the page
What is the office wifi password? Not in the docs The server says it is not covered
  1. Ask each row in a fresh chat with the server enabled. Do not hint at the page title.
  2. Use the client people actually use. A second client is a second run, labeled separately.
  3. Save the date, the client, and the fails.

The dashboard shows that the calls happened. It does not compare a fact to a page. That comparison is the sheet. Action analytics shows the passage so the comparison is faster, which is Measure MCP retrieval quality.

Check that it works

A pass on the upload header is the right page and the right header name. A pass on the wifi question is an explicit miss. An early run and a finished run are not comparable. Say so if you keep the early numbers.

If a row fails

Fix one fail using returning the wrong information, then rerun the whole list. If two pages could both be correct, pick the canonical one in the sheet. If the other keeps winning, prefer the canonical source with a retrieval rule. Headings that match the question are Make your documentation AI-ready.

Write the sheet and run it today

Open your server in MCP Studio after indexing shows complete. Ask the three questions in the table. Keep a row only when the link and the fact both match, or when the out-of-scope question is refused.

Where to go next