2026-10-01 • 5 min read • Build
What a Context Engineer Puts in Front of the Model
An LLM writes from the context that arrives with the question. A context engineer decides which slice of the stack you already have shows up at that moment, and keeps it available through an MCP server.
An LLM writes from the text that arrives with the question. It does not open your docs site, your repository, or last quarter's incident write-up on its own. If a page is not in that request, it does not inform the output, however carefully someone wrote it.
Most of that material already exists. The work of a context engineer is to decide which slice of it a given moment needs, and to make that slice reachable every time the moment comes around.
This piece is about that job. The companion piece looks at the same servers from the shape of the system: the context a software architect keeps with the stack.
The stack is already full of context
A normal product stack holds four kinds of answers, each in a system that is good at its own job.
- Public docs, pricing, and the changelog say what you promise.
- The repository says how a change is made here.
- Design reviews and decision records say why a rule exists.
- Runbooks and postmortems say what happened last time.
A person moves between these by remembering where to look. An LLM receives one bundle of text, writes from it, and does not keep the bundle when the chat ends. The pages can be excellent and still never meet the question.
The moment decides the slice
The same team, on the same day, needs different pages depending on what they are doing.
| Moment | The question in front of the model | Where the answer already lives |
|---|---|---|
| Planning a feature | What do we promise? | Docs, pricing, changelog |
| Writing the change | How is it done here? | The repository |
| Reviewing the change | Why is it this way? | Decision records, design reviews |
| Running it | What happened last time? | Runbooks, postmortems |
Hand all four over on every question and the useful passage is crowded out by the other three. A context engineer names the moment first, then gives that moment its own MCP server: one address, one kind of question, sources that belong together.
When the server is there, the model writes from a passage of yours, and the passage names the page it came from. When it is not, the model writes from general knowledge, and the reply has nothing in your stack to check it against.
That citation is the part a reader can use. Anyone can open the page and see whether the reply followed it.
Persistent means the next tool asks the same place
Pasting a page into a prompt works for one chat. The next person, in another client, pastes a different page, or none, and the model is working from a different stack than the one you maintain.
An MCP server is a stable address for one slice. Cursor, Claude, Codex, and whatever you connect next can all call it. The chat is temporary. The server is still there tomorrow, with the same sources, so the context that informed yesterday's answer is the context available today. When the docs change, you refresh the source, and the next request sees the new page. You do not have to find every prompt someone saved.
That is the sense in which context is persistent. The model does not grow a memory of your company. The stack keeps an address the model can ask, at the moment the question shows up.
You can see whether the slice arrived
A server that exists is not yet a server that was used. Every request is recorded. On any account, the dashboard shows how often the server is asked and how that changes over time. The Action analytics add the passage that came back, and they mark requests that came back empty or only weakly matched. The analytics guide describes each view.
An empty or weak result is a moment you have not covered yet. Either the page is missing, or it exists and is written so the question cannot find it. Both are edits to the sources, which is the part of the stack a context engineer can change directly. Read that list, fix the page, refresh the source, and ask the question again.
Giving a moment its address
MCP Studio turns sources you already have into that address. You add a documentation URL, a GitHub repository, an uploaded file, or another MCP server, choose the tools, and deploy. The source guide covers each kind. You get one endpoint any MCP client can connect to, and each answer cites the page it came from.
Pages that should stay inside the company go on a private server, which requires an access token before it will answer, including when a client asks what tools it has.
What this does not do
The server supplies the passage. The model still writes the sentence. If the passage is not in your sources, the useful result is an empty reply that tells you the page is still unwritten, rather than a fluent guess. One server that holds every system at once tends to return a blend, which is why the split by moment is the design.
Companies draw those moments differently. The four rows above are a shape that fits a lot of product teams, not a template to copy line for line. Start from the question your own tools keep answering without your pages.
Pick that question and give it one server. MCP Studio takes a docs URL, a repository, or a file and returns an endpoint you can connect.