← Back to Blog

2026-10-01 • 5 min read • Build

The Context a Software Architect Keeps With the Stack

A software architect's decisions inform an LLM only when they are available at the moment of the change. MCP servers keep that context beside each layer of the stack you already run.

A software architect already draws the stack: services, boundaries, and the reasons a boundary exists. Payment writes carry an idempotency key. One service owns refunds. A public API does not expose a field the changelog has not announced.

An LLM working in that stack sees the files in front of it and the text in the prompt. A decision informs the output when it is part of that text. Otherwise the model produces something plausible that may not be the decision you recorded.

The architect's addition is not a second stack. It is an address on each layer of the one you have, so the decision is fetched at the moment a change is planned, written, reviewed, or run.

This piece is about those addresses. The companion piece is about the pages inside them: what a context engineer puts in front of the model.

The decisions are already written down

They live where your team put them. A decision record. A design review. A comment in the repository that is actually the rule. A runbook from the last time this failed.

People on the team carry a map of those places. A tool that generates code does not. It will follow a pattern it has seen in other codebases unless the pattern from this codebase is in the request. The architect's map has to be something the tool can call, or the map stays in people's heads while the output comes from somewhere else.

Each stage of a change needs a different layer

A change crosses several moments, and they do not want the same pages.

  • Planning needs what you promise, so the design does not invent a behavior you do not offer.
  • Building needs how this codebase does the thing, so the change matches the code beside it.
  • Reviewing needs why it is this way, so a rule from a design review is present when the diff is read.
  • Operating needs what happened before, so a fix is read against the last incident.

Put those in one pile and every stage receives pages it will not use. Split them into MCP servers, one job each, and the model at that stage receives the layer that governs it.

Four MCP servers, each with one job, feeding the stage of a change that needs it. Product feeds planning. Decisions feed planning, building, and review. The codebase feeds building and production. Operations feeds production.

The decisions server shows up in three stages. That is a hint about where an architect's attention goes: the constraints that apply in more than one place are the ones most often missing from a single prompt.

The shape is yours to draw. Some stacks want a server per service, or per domain, rather than the four above. The useful test is whether a person at that stage can name the question the model should be able to ask.

One address, every tool in the flow

A prompt is a private copy. Two people, in two clients, can hand the model two different versions of the same rule, and both replies will sound confident.

A server is shared. The person planning in Claude, the person building in Cursor, and a check you run in a pipeline can all call the same endpoint, if you point them at it. When the decision changes, you update the source once. The next request, from any of those tools, reads the updated page.

That is what persistent context means across a stack. Not a longer memory inside the model, and not a prompt someone has to remember to paste. A stable place, beside the layer the decision belongs to, that any LLM in the flow can ask at the time it acts on that layer.

What changes in the output

The model still writes the code and the explanation. What changes is the material those are written from, and whether a reviewer can see it.

A reply that cites the decision record can be checked in a minute: open the page, compare. A reply built from general habit can be checked only by someone who remembers the decision and thinks to look. Citations keep the stack visible inside the output, instead of letting it disappear into a fluent paragraph.

You can also see which layer the tools are actually consulting. Usage over time shows whether a server is being asked. A decisions server that nobody calls is a constraint the flow is not reading, which is worth knowing before you add another one. The companion piece covers how to read empty and weakly matched requests, which tell you the page is missing or hard to find.

Putting an address on a layer

MCP Studio turns the sources for one layer into an endpoint. You add a documentation URL, a repository, an uploaded file, or another MCP server, choose the tools, and deploy. Each answer cites the page it came from. The source guide covers the kinds of source.

Architecture decisions are often internal. A private server requires an access token on every request, so the constraint is available to your tools and not to the open internet.

What this does not do

The server returns what you placed in it. A decision that lives only in a meeting does not inform the model, and no arrangement of servers fixes that. The model can also be given a passage and still not follow it. The citation is how a reviewer sees the difference.

Start with the constraint your reviews keep repeating, and give that layer an address. MCP Studio takes a docs URL, a repository, or a file and returns an endpoint the tools in your flow can share.