2026-09-22 • 5 min read • Measure
How to Measure MCP Retrieval Quality
How to score whether the cited passage contains the fact, on a fixed list you can rerun in MCP Studio after a docs edit.
A server can be busy and still hand the agent the wrong paragraph. Request volume, on the free dashboard, tells you the client called the server. It does not tell you the passage contained the fact. Retrieval quality is that second question, and you answer it by reading.
What you are scoring
The passage, not the prose wrapped around it. For each question, a few words of fact. "Header is X." "Overlap is 24 hours." The same questions, the same expected facts, rerun after you edit. A list that changes every week cannot show whether the edit helped.
The yes-or-no version of the list is Test an MCP server for accuracy. A heavier way to mark an answer is the MCP server rubric.
Why volume is not quality
Yes if the cited passage contains the fact.
No if the passage is a different section, a different source, or the right page without the fact. If the constraint the person needs is missing, that is a no, even when the topic is right. Why a loosely related paragraph still produces a confident answer is Choosing the right context.
The out-of-scope question is a yes only when the server says the content does not cover it. A polished sentence can restate the wrong page. You will give it credit if you only read the reply.
Read the passage in MCP Studio
MCP Studio is the no-code builder on Appa Tools. Open the server on the dashboard after indexing is complete.
- Ask each question once, with the server enabled.
- Record the cited page and copy the passage. Action analytics, $99 a month with Core, shows the page, section, and excerpt so you are not copying passages out of the chat by hand. Core at $49 a month shows the question text. The free tier shows source and tool. The table is in What the questions people ask your docs can tell you.
- Mark yes or no against the fact, not against how smooth the reply sounds.
- For each no, one edit. A heading, a source, or a retrieval rule. Then rerun the whole list. Do not average a partial rerun into the old count.
Which pages those passages came from, as a habit rather than a single test, is Measure which documentation an AI agent actually uses.
Check that it works
Write the result as a record of one run, with the limits attached. An example of that sentence, not a product score: "4 of 5 answerable questions returned a passage that contained the fact, on the Billing API server, in Claude, on September 22, 2026, after indexing completed. The fifth cited the overview. The list is ours."
That is not a universal score, and it is not comparable to a three-question run on a different server.
If a passage is a no
Move the missing constraint next to the steps, or prefer the right source. Then rerun the full list. The edit path is Fix an MCP server returning the wrong information.
Score one fixed list today
Open the server in MCP Studio and ask the same five questions you wrote down. Mark each passage yes or no. Keep the count next to the server name, the client, and the date.