2026-08-21 • 9 min read • Measure
Confident, Well Written, and Wrong: A Rubric for Docs MCP Accuracy
An answer can be fluent and still wrong in the one detail that matters. Here is a repeatable way to score a documentation MCP server across five dimensions, with a worked example and a full test you can run yourself.
You ask an AI assistant how to move Metabase off its embedded database and onto a production one. Back comes a clean answer: the command, a flag, a sentence explaining what the flag does. The command is correct. The flag does not exist.
That is the failure worth building a test around. Not a garbled reply, but a confident, well written paragraph with one wrong detail in the middle of it and nothing in the wording to tell you which one.
Connecting a documentation MCP server to your client is meant to fix this. Whether it does is something you can settle with evidence, and this post is the method. It works on any MCP server, ours or anyone else's, and needs no account.
"Is it accurate" is not one question
Ask whether an answer is accurate and you get a yes or a no that hides most of what you need to know.
An answer can be true and too vague to act on. It can be precise and describe a release from two years ago. It can be right on every fact and still omit the prerequisite that makes the procedure work. It can be correct throughout and impossible to trace to a page, which leaves you no way to check it.
Different problems, different repairs. Score them separately and the result tells you what to fix. That is all the MCP Server Accuracy Rubric is: five dimensions, ten points each, so one question is worth 50 and a five question test is worth 250.
Write the answer down before you ask the question
Write out the answer you expect first, with the page and section supporting each fact in it. A fluent answer is persuasive, and if you decide what counts as correct while reading one, it will talk you into crediting details you never verified.
Start with facts that have a single defensible answer. Version requirements, exact commands and flags, environment variable names, accepted parameter values, migration steps, documented error messages. Add conceptual questions once everyone who scores agrees on what a good one looks like.
The five things worth scoring
Carry one real question through all five. Ours is question four of the test below: what command migrates Metabase data from H2 to a production database, and what is the important detail about the file path?
Factual accuracy asks whether every checkable claim matches the current
documentation. The unconnected answer to our migration question named a
--h2-jar-path flag that appears nowhere in the documented command. A polished
explanation does not buy back a wrong value.
Specificity asks whether you were handed the exact values you need to act.
MB_JETTY_PORT is actionable. "Change the Jetty port setting" points the right
way and leaves the work with you.
Completeness asks whether the answer covers every part of the question, plus the caveats you need to use it. Our migration question has two parts, and an answer that nails the command while skipping the file path detail fails you at the point it matters.
Currency asks whether the answer describes the current release. It earns a dimension of its own because an answer can faithfully describe what was true two releases ago and be useless today, and the phrasing will not warn you.
Groundability asks whether you can trace each claim to a specific statement. Go and look. When a claim is correct and you still cannot find its source, note it, because sometimes your own page states the fact too indirectly for anyone to find.
Score each dimension out of ten on the same scale.
| Score | What it means |
|---|---|
| 10 | Nothing to correct on this dimension |
| 7 to 9 | One minor detail off, loosely paraphrased, or hedged |
| 4 to 6 | The core is right, and one specific piece is wrong or missing |
| 1 to 3 | Pointed the right way, with the important values wrong |
| 0 | Substantively incorrect, vague, or invented throughout |
Score a question your documentation cannot answer
Here is the part we would argue for hardest.
Put at least one question in your set that your content genuinely does not cover, and score it on its own. A clear "this is not in the documentation" is full marks. A fluent, plausible, unsupported answer is zero, however well it reads.
That feels backwards the first time you write it down, because you are giving a server its best score for producing less. It is the right way round. Every other question measures how good an answer can be. This one measures whether you are entitled to believe any of them, because a server that always produces something would have produced those too.
You: How do I set up automatic backups of my dashboards?
Assistant: Nothing in the connected documentation covers automatic dashboard backups. The closest material is about backing up the application database, which is a different thing.
That is not an impressive answer. It is an honest one, and it is the difference between a tool you can act on and a tool you have to check behind. If every question you test is one your documentation answers well, this is the single behavior your test cannot see.
Running the comparison fairly
Run the same questions twice, once with the documentation server available to the client and once without, and hold everything else still:
- Model and version
- System instructions
- Prompt wording
- Question order
- Sampling and reasoning settings
- Client and session configuration
Say in the prompt what you expect of the connected run. "Use the connected documentation tools and cite the sources you rely on" is enough. Save the answers, the tool calls and the passages that came back, because the passages are what you score groundability against. Better still, have a second person score them without knowing which run is which.
A worked test on Metabase's documentation
We ran this in August 2026 against Metabase documentation, tutorials, use cases, and its public documentation repository. The five questions:
- What Java version is required to run Metabase?
- What environment variable changes the port Metabase runs on from its default?
- What are the exact environment variables needed to connect Metabase to a Postgres application database?
- What command migrates Metabase data from H2 to a production database, and what is the important file path detail?
- What is the minimum supported MySQL version for the Metabase application database?
The expected answers were written first, from these pages:
- Java version, running the Metabase JAR
- Port variable, environment variables
- Postgres and MySQL application database details, configuring the application database
- Migration command, migrating from H2
- Migration error, troubleshooting load from H2
Those pages were the ground truth in August 2026. The links point at the latest Metabase documentation, so what is on them now may have moved on.
The scores
Without the documentation server connected:
| Question | Factual | Specificity | Completeness | Currency | Groundability | Total |
|---|---|---|---|---|---|---|
| Q1, Java version | 2 | 6 | 6 | 1 | 3 | 18 |
| Q2, port variable | 10 | 8 | 7 | 10 | 7 | 42 |
| Q3, Postgres variables | 9 | 9 | 6 | 9 | 6 | 39 |
| Q4, H2 migration | 4 | 5 | 5 | 6 | 3 | 23 |
| Q5, MySQL minimum | 2 | 6 | 4 | 2 | 3 | 17 |
| Total | 27 | 34 | 28 | 28 | 22 | 139 / 250 |
With it connected:
| Question | Factual | Specificity | Completeness | Currency | Groundability | Total |
|---|---|---|---|---|---|---|
| Q1, Java version | 10 | 10 | 9 | 10 | 9 | 48 |
| Q2, port variable | 10 | 10 | 9 | 10 | 9 | 48 |
| Q3, Postgres variables | 10 | 10 | 10 | 10 | 9 | 49 |
| Q4, H2 migration | 10 | 10 | 10 | 10 | 10 | 50 |
| Q5, MySQL minimum | 10 | 10 | 10 | 10 | 10 | 50 |
| Total | 50 | 50 | 48 | 50 | 47 | 245 / 250 |
The migration question we carried through is the clearest single case: 23 points without the documentation server, 50 with it, and no invented flag.
Read those totals for what they are: five questions, one setup, on documentation we do not own, scored by us. The questions lean towards version specific and syntax exact facts, which is the kind of question grounding helps most, and the rubric asks a person to judge. Treat this as one controlled comparison rather than a benchmark, and run your own.
What the question by question differences tell you
A stable fact gains the least. The port variable question scored close in
both runs, because the unconnected answer already knew MB_JETTY_PORT.
Grounding tightened the wording and added a citation.
Version facts are where the gap opens. The unconnected answer to the Java question said Java 21, while the documentation at test time required Java 25. Nothing in the wording marked the difference, which is why currency is scored on its own.
Invented specifics are the reason groundability is scored at all. That
--h2-jar-path flag reads like a real one. The connected answer gave the
documented command, the requirement to leave .mv.db off the path, the version
requirement, the empty target database requirement, and the documented
troubleshooting message, drawn from more than one page.
Watch for answers borrowing from an earlier question. The MySQL question never triggered a lookup of its own. A passage fetched for the migration question happened to list supported application database versions, and the answer came from there. It scored well, and that score is less repeatable than it looks, which is what asking every question in one prompt costs you. Ask each one in its own session if you want to know.
Turn the lost points into a maintenance loop
Every dimension names its own repair. Factual losses mean a source is wrong or two pages disagree. Specificity losses mean an exact string or command is missing from the page. Completeness losses mean a prerequisite sits too far from the procedure. Currency losses point at archived pages and missing version labels. Groundability losses mean a fact is true on your site and stated too indirectly to find.
Keep the questions stable so the scores stay comparable, add new ones for each release and repeated support topic, and rerun with the same model, prompts and settings after any change to your content.
The MCP Studio analytics dashboard shows which pages were retrieved, which helps when you score groundability. If you need a connected condition to test, Context MCP Studio can put documentation sites, GitHub repositories, PDFs and other MCP sources behind one endpoint, and the free tier covers a first server.
Download the MCP Server Accuracy Rubric (PDF)
Keep going
- Create your first MCP server, a practical starting point
- Connect to Claude, Cursor, or VS Code, setup for the connected condition
- How MCP Studio works, from source to retrieved answer
- Docs MCP metrics, signals you can review after launch