← Back to Blog

2026-10-01 • 14 min read • Build

Playbook: From Technical Writer to Context Engineer, From Engineer to Architect

A step-by-step playbook for using AI, MCP Studio, and a Jev decision model to move from technical writer to context engineer or from engineer to architect, with diagrams, a 90-day plan, and the analytics that show your impact.

We are all capable of more than ever with AI. Drafting, coding, and searching take a fraction of the time they did. That raises the question this playbook is about: how do we use the extra capacity to level up, and not just to produce more of the same?

For two common starting points, the answer is a role that is still being named. A technical writer can become a context engineer: the person who decides what an AI system can know, shapes that material so the right passage comes back, and proves it works. A full-stack engineer can become a software architect: the person who owns the decisions and constraints every change must satisfy, and builds the checks that enforce them.

The two roles meet in one place, a knowledge architecture made of focused MCP servers that agents, people, and automated checks all read from. By the end of this playbook you will have a map of those servers, a first one built in MCP Studio, a scored set of test questions, a decision-model check on your agent's work, and a dashboard that shows whether any of it is helping.

What context engineers and architects do

Titles vary by company, so treat these as working definitions for this playbook.

A context engineer owns what an AI system can know and how reliably it finds it. The daily work:

  • Decide which questions the AI will face and which source answers each one.
  • Shape sources so one passage answers one question.
  • Keep a set of test questions with known answers, and run it after every change.
  • Read usage data to find questions the sources do not answer, and close the gap.
  • Write the criteria that automated checks use to judge work.

A software architect owns the decisions and constraints that every change must satisfy, and the way they are enforced. The daily work:

  • Record decisions with their reasons, in a place a tool can read.
  • Define boundaries: what is checked by code, what by a model, what by a person.
  • Turn team rules into questions a check can answer.
  • Set thresholds and escalation paths, and review the exceptions.
  • Run the feedback loop that keeps all of it current.

Read the two lists together and the overlap is plain. Both roles design how knowledge and judgment move through a system. One approaches it from the content side, the other from the system side, and a team needs both.

Where MCP Studio fits

AI clients such as Cursor, Claude Desktop, and Windsurf reach outside information through MCP, the protocol they use to call tools. The material that matters to your team lives in documentation sites, repositories, PDFs, spreadsheets, and other MCP servers. The question is how to turn that material into something an agent can query.

MCP Studio turns the sources you already have into hosted MCP servers without code. Paste a documentation URL, connect a GitHub repository (private ones included), upload a file, or add another MCP server. Choose the tools you want and deploy. Every server gets one endpoint any MCP client can connect to, and every answer carries citations to the pages it came from.

That matters for this playbook because it moves the work to where a writer or architect has the most leverage. The effort goes into which servers exist, what goes into each, how they are tested, and how their impact is measured. The server plumbing is handled. Along the way you also get private servers with access tokens, retrieval rules, and analytics, each of which has a step below.

Step 1: List the questions your work depends on

Before you build anything, write down the questions that come up repeatedly in your team's work. Aim for 15 to 20 and sort them by the kind of answer they need:

  • What do we promise customers? Product behavior, limits, pricing, policies.
  • Why is it built this way? Decisions, trade-offs, constraints.
  • How is it done here? Code conventions, patterns, examples.
  • What has happened before? Incidents, runbooks, past fixes.

If you have never done this, ask your AI assistant to help you mine it. The assistant proposes, and you decide:

Here are our last 30 pull request review comments [paste them]. Group them by
the kind of knowledge the reviewer relied on. For each group, name a question
a coding agent should be able to answer before it writes code. Do not invent
groups the comments do not support.

The last sentence of that prompt is the habit that defines this role. You use AI to widen what you can see, then you check the result against evidence.

Step 2: Give each job its own server

One big server that holds everything returns a mix of everything. Separate servers, each with a narrow job, return passages that belong together. A common split looks like this:

Server Answers Typical sources
Product What do we promise customers? Public docs site, pricing page, changelog
Decisions Why is it built this way? Architecture decision records, design review PDFs, specs
Codebase How is it done here? The repository, with its docs folder
Operations What has happened before? Runbooks, postmortems, support guides

A map of four MCP servers, Product, Decisions, Codebase, and Operations, each feeding the stage of work that needs it: planning, building, checking, or running in production.

Each server feeds the stages that need it. The Decisions server appears in three places, which is a signal that it deserves the most care.

The same map as text, so you can paste it into your own notes or diagram tool:

flowchart LR
  subgraph servers [Four servers, one job each]
    P["Product<br/>What do we promise?"]
    D["Decisions<br/>Why is it this way?"]
    C["Codebase<br/>How is it done here?"]
    O["Operations<br/>What has happened before?"]
  end

  P --> Plan[Plan the change]
  D --> Plan
  C --> Build[Build the change]
  D --> Build
  D --> Review[Check the change]
  O --> Operate[Run it in production]
  C --> Operate

Software architects: this map is a design artifact. It records which knowledge exists, who owns it, and which stage of work depends on it.

Context engineers: this map is your scope. Each box is a body of material you are responsible for.

Step 3: Build your first server in MCP Studio

Start with the server whose absence hurts most. For most teams that is the Decisions server, since decisions are written down least consistently and needed most often.

  1. Open the wizard at MCP Studio and add your first source. A documentation URL, a GitHub repository, an uploaded file, or an existing MCP server all work. The source guide explains each type, and the upload guide covers PDFs, spreadsheets, Word files, slides, and images.
  2. Choose tools that match your questions. The tool guide explains what each one is for.
  3. Choose visibility. Decisions and design reviews are usually internal, so make the server private. A private server requires an access token on every request, including listing its tools.
  4. Deploy, then connect a client. The guides for Cursor and Claude walk through the connection.
  5. Ask the server three questions you already know the answer to. If it cites the right page, you have a working server.

Free accounts include 50 MCP requests a month, shared across all your servers, so run your questions deliberately. Paid request plans are on the pricing page.

Step 4: Shape the sources for retrieval

This is the most transferable skill for a technical writer, and the one an architect should learn. Retrieval returns passages, so the unit of knowledge is the passage, not the document.

  • One decision per page, with the rule in the title. "Payment writes must carry an idempotency key" tells a reader and a search system the same thing. "Payments design review, March" tells neither.
  • State the rule, the reason, and the exception in that order. A passage that starts with background may be cut off before it reaches the rule.
  • Use the words your team uses. If people say "invoice", the page should not only say "billing document".
  • Keep headings specific. Headings anchor the sections a citation points to.
  • Keep one source of truth per fact. Two pages that disagree produce two confident, conflicting answers.

You do not have to rewrite everything. Start with the 10 pages your question list points to most often. AI is useful here as an editor: ask it to rewrite a page so each section answers one question, then read the result against the original to confirm it kept every exception.

Step 5: Build and score a question set

A server that looks right is not the same as a server that is right. A question set is how you tell them apart, and it is the artifact that best shows a context engineer's work.

Take the questions from Step 1 and record, for each, the page that should answer it:

Question Page that should answer it Cited the right page? Passage enough to act on?
Do payment writes need an idempotency key? Payment write rules Yes Yes
Which service owns refunds? Service ownership No n/a

Run the whole set, mark each row, and keep the score. Two columns matter. The first checks that the right page came back. The second checks that the passage was enough to act on, because a page that is cited but cuts off mid-instruction is not a success. Write one number for each column.

Re-run the set after every change to a page, a source, or a rule. If the number falls, the change did harm, and you know which change it was. For a longer treatment of testing, see how to test an MCP server for accuracy.

Step 6: Teach the server what search cannot know

Some answers are not findable from the question. A brand kit image shares no words with "write a customer welcome email", and no ranking can connect the two. A retrieval rule is a plain-sentence instruction that tells the server what to prefer or include for questions of that kind.

Rules are the place where an architect's judgment becomes part of the server. Write them for the cases your question set shows are missed, and keep them narrow. A rule steers retrieval. It does not change what your content says.

Step 7: Put the servers in the workflow

With servers in place, decide who does what at each stage of a change. A useful division of labor has four parties:

  • Code does exact work: permissions, arithmetic, counting, dates, and anything with a precise right answer.
  • A decision model makes bounded judgments, such as whether a change follows a rule.
  • An LLM agent reads, plans, and writes.
  • A person handles the risky and the uncertain.

A workflow in which an agent plans and builds a change using the Product, Decisions, and Codebase servers, a decision model checks the work, and a person reviews anything uncertain.

The agent drafts. A check judges the draft against the Decisions server. Code enforces what is exact. People see only what the checks could not settle. The design question for an architect is where each boundary sits, and the question for a context engineer is whether the check has the right text to judge against.

Step 8: Add a decision-model check

Most of this workflow works with any agent. The check in this step uses Jev, a decision model from TypeSafe. Jev does not write prose. You give it a small piece of state and a typed question, and it returns probabilities. As described in OpenRouter's tutorial on when to use Jev, the question types are choice (pick one option), noul (the probability of yes), and score (a number on a scale).

That shape suits a rule check, because "does this change follow this rule?" is a bounded question with a small set of answers. It also comes with limits that shape how you build:

  • It reads criteria literally, so write them like a specification.
  • Its accuracy drops when you include state the question does not need, so send only the relevant passage and the relevant part of the change.
  • It is not dependable at arithmetic, counting, or dates, so leave those to code.
  • Ask one small judgment per question.

This is the point where an MCP server contributes most. The server supplies the focused state: the one passage from the Decisions server that governs this change. The architect supplies the question. The shape of a closed question, written as pseudocode rather than any specific API:

state:
  rule:    <passage returned by the Decisions server>
  change:  <the part of the diff that touches payment writes>

question: Does the change follow the rule?
options:
  supported      - the change sets an idempotency key on every payment write
  unsupported    - at least one payment write has no idempotency key
  not_addressed  - the rule does not apply to this change

Each option is defined by what would make it true. Vague options such as "mostly fine" give the model nothing to read literally.

A sequence in which the agent proposes a change, the Decisions server returns the governing passage, the decision model returns probabilities for supported, unsupported, and not addressed, and code decides whether to pass, block, or send the change to a person.

Code makes the final call. If the probability of "unsupported" is high, block the change. If the model is unsure, or the answer is "not addressed" on a change that clearly should be covered, send it to a person. The thresholds are your policy, so tune them on your own labeled examples instead of copying a number. A practical way to start is shadow mode: run the check for two weeks, record its verdicts next to what your reviewers decided, and set thresholds only once you can see where they agree.

Jev's decisions endpoint is in alpha at the time of writing, so check OpenRouter's tutorial for the current interface before you wire it in.

What an MCP server contributes to any workflow

Everything above uses a decision model, but the server's value does not depend on one. In any workflow, a well-built MCP server does four things:

  1. It gives the work an address for knowledge. A decision that lived in a chat thread now has a place to be asked.
  2. It supplies focused context. Agents and checks work better with the one relevant passage than with a whole wiki, and a decision model needs exactly that.
  3. It cites its sources. Every answer points back to a page, so a person can verify it in seconds.
  4. It records what was asked. Every request to the server is logged, whether it came from an agent, a teammate, or your own check. That log is the raw material for measuring impact.

Step 9: Read your impact in the analytics

An architecture you cannot measure is an opinion. MCP Studio's dashboard turns usage into evidence, in layers that build on each other.

Layer What you see What it tells you
Free for every account KPI cards, Calls Over Time, and a Breakdown Whether the server is used, how often, and by what
Core A request log The questions people and agents actually asked
Action The exact passage returned for each request, retrieval confidence, content gaps, and an impact model Whether the right content is being used, and where it is missing
Predictive Requests to a working answer, source grades from A to D, ranked fixes, and a 30-day forecast Which sources to improve first, and what load is coming

New accounts get a 30-day trial with all of the analytics. The analytics guide and dashboard guide describe each view.

The Action tab is where a context engineer spends time. Its summary strip shows the share of requests answered from your content, how many came back empty, and how many matched only weakly. An empty or weak result is a question your sources do not answer: either the page is missing, or it exists and is written so the question cannot find it. Both are fixes you can make.

The Predictive tab adds the measure that matters most for a Jev workflow: requests to a working answer. If an agent or check needs five requests to reach a usable passage, the content is not shaped for the question. Source grades from A to D point you to the weakest source first.

The impact model shows its formula. Dollar figures are opt-in, because a monetary result depends on assumptions about what your time costs, so nothing is priced until you choose how.

A loop in which agents and checks ask the server, the server answers with cited passages, the dashboard shows what was used, matched weakly, or came back empty, you fix a page or add a rule, and you refresh the source and re-run your question set.

Connecting this to the Jev check. Because your check queries the server like any other client, its requests appear in the same dashboard. A check that keeps returning "not addressed" on changes that should be covered tells you the Decisions server is missing a page. A check with low retrieval confidence tells you a page exists but is not shaped for the question. Either way, the fix lands in the content, which is where a context engineer has the most leverage.

A weekly routine that fits in half an hour:

  1. Open the Action tab and read the summary strip.
  2. List the questions that came back empty or weak.
  3. For each one, decide: write the page, rewrite the page, add a retrieval rule, or accept that it is out of scope.
  4. Refresh the source and re-run your question set from Step 5.
  5. Write down the change in each number.

Using AI as your coach

You can use AI at every step, as long as you keep the final judgment. Four prompts that work well:

Here is a page from our wiki [paste]. Rewrite it so each section answers one
question and the rule comes first. List every exception from the original that
you kept, so I can check nothing was lost.
Here are 20 questions our agents should be able to answer [paste]. For each,
write the page title that should answer it. Flag any question where you would
need information I have not given you.
Here is a rule from our architecture decisions [paste]. Write three options for
a closed question that checks a code change against it. Define each option by
what would make it true, using only facts visible in a diff.
Here are this week's empty and weak questions from our dashboard [paste]. Group
them by cause: missing page, badly shaped page, or out of scope. Do not guess
at causes the questions do not support.

Each prompt asks the model to show its work, so you can verify it. That is the skill in a nutshell: use AI for range, and use evidence for trust.

Your 90-day path

The two tracks differ in where they start and converge on the same artifacts.

Two career paths, technical writer to context engineer and full-stack engineer to software architect, both leading to a shared knowledge architecture of MCP servers, a question set, and decision checks.

When Technical writer track Engineer track Artifact you can show
Weeks 1 to 2 Collect questions from support threads and reviews (Step 1) Collect recurring review comments (Step 1) A question list sorted by job
Weeks 3 to 4 Rewrite the 10 most-needed pages for retrieval (Step 4) Build the first server, private (Step 3) A working server and an edited page set
Weeks 5 to 6 Build the question set and score it (Step 5) Turn two team rules into closed questions (Step 8) A scored question set, a rule-to-question list
Weeks 7 to 8 Add retrieval rules for missed cases (Step 6) Run the check in shadow mode (Step 8) A rule log, a shadow-mode comparison
Weeks 9 to 12 Run the weekly analytics routine (Step 9) Set thresholds from the comparison, add escalation Before and after numbers with their method

What each track already brings:

If you are a technical writer What you bring What you add
Structure and clarity Pages shaped for retrieval Test sets and measurement
Knowing what readers ask The question list Usage data to confirm it
Maintaining accuracy One source of truth per fact Automated checks that depend on it
If you are an engineer What you bring What you add
Systems thinking Boundaries between code, models, and people Knowledge as a first-class part of the system
Rules and review habits The constraints worth checking Closed questions a check can judge
Shipping Working integrations Measurement of whether context helps

Keep the evidence. A before-and-after figure from your own dashboard, such as the share of requests answered from your content or the average requests to a working answer, is a credible portfolio item when it travels with its method: what you changed, over what period, and on how many questions.

Limits worth knowing

  • A question set measures the questions you wrote. Keep adding new ones from real usage so it stays honest.
  • A decision model gives probabilities, not proof. Keep a person on the cases the threshold cannot settle.
  • A server returns what your sources contain. It can point to a gap, but it cannot fill one, so the pages remain your responsibility.
  • Titles and expectations differ between companies, and a new title does not follow automatically. The artifacts are what make the change visible.

Frequently asked questions

What is a context engineer?

A context engineer decides what an AI system can know and makes sure it finds the right information reliably. The work includes choosing sources, shaping them for retrieval, testing them with known questions, and reading usage data to fix gaps.

How does a technical writer become a context engineer?

Start from the skills you already have: structure, clarity, and knowing what readers ask. Add three habits: shape pages so one passage answers one question, keep a scored question set, and use analytics to find the questions your sources miss. The 90-day path above lays out the order.

How does a full-stack engineer become a software architect?

Move from writing changes to owning the constraints every change must meet. Record decisions where tools can read them, turn rules into closed questions a check can answer, and decide which judgments belong to code, a model, or a person. The artifacts are a decisions server, a rule-to-question list, and a shadow-mode comparison.

Do I need to write code to build MCP servers with MCP Studio?

No. You add sources in a wizard, choose tools, and deploy. The server gets a single endpoint you connect to a client such as Cursor or Claude Desktop. You only write code if you want to call the server from your own automation, such as a check in your workflow.

What does Jev do in this workflow?

Jev is a decision model that returns probabilities for typed questions, such as choosing between options or estimating the chance of yes. In this playbook it judges whether a change follows a rule, using a passage from your Decisions server as its input. Code and people handle what a probability cannot settle.

How do I show that my MCP server is making a difference?

Use the dashboard. Free accounts see usage and trends. The Action tab shows the share of requests answered from your content and how many came back empty or weak, and the Predictive tab shows requests to a working answer. Record those numbers before and after each change, with the date range and number of questions.

Does an MCP server help if I do not use a decision model?

Yes. A server gives your knowledge an address, supplies focused context to any agent or teammate, cites its sources, and records what was asked. The decision model is one consumer among several.

Where to go next

Pick one question from Step 1 that your team keeps asking, and build the server for it. MCP Studio takes a documentation URL, a repository, or a file and returns a connectable endpoint. The first question set takes an afternoon, and the first week of analytics will show you what to fix.