Education

How to design an AI agent in 2026: the system design, block by block

If you built an AI agent from scratch today: the 10 blocks to attach, what goes inside each one, and why the data layer decides whether the rest works.

How to design an AI agent in 2026: the system design, block by block

If I were building an AI agent from scratch today, I'd make ten decisions before writing a single prompt. Here's the whole system on one page, then each block you attach: what goes inside it, the steps to add it, and the trick that saves you a week.

Try GuidedMind.ai free → The full AI agent system

The whole system. Numbers match the ten blocks below. The plum block is where most agents quietly fail.

Read it like a build list

Every section below is one block you attach to the system. Each one has the same five parts, so you can skim to what you need:

  • What's inside: The parts the block is made of, as a checklist.
  • Steps to attach it: The order to add those parts in, without rework.
  • The trick: The one thing that saves the most time later.
  • Go deeper: The design details that matter once it's live.
  • Read more: The research papers each block is built on.

💻 Get the code. Every block in this article is a small Python module in our open-source boilerplate: context, decisions, tools, memory, harness, loop, guardrails, evals and tracing, wired into a working agent with offline tests. Boilerplate on GitHub →

Half of the ten blocks (context, memory, tools, evals and observability) are really about one question: what does your agent actually know, and how do those facts connect? That's why the context layer comes first and gets the most space.

1 · Context layer: don't build your own RAG pipeline

Attach first · the data layer

"RAG" isn't one technique anymore. Plain vector search, hybrid search with keywords, GraphRAG, agentic retrieval, corrective retrieval, multimodal. The right one depends on the questions your users ask, and you will want to switch as you learn. Hand-building one pipeline locks you into the first guess and buries you in a data-layer rabbit hole: chunking, embeddings, vector databases, re-indexing.

So instead of a pipeline, attach a context layer: one place where your documents are processed, turned into a knowledge graph, retrieved the right way per use case, and tested and traced before your agent depends on them.

Documents become chunks and a knowledge graph; a question is answered by following the relations Mia to curly cut to Saturday, then scored

Vector search finds paragraphs that look similar. A graph follows how facts connect, and a score tells you if it worked.

What's inside

  • Connected sources: docs, sheets, websites
  • Processing per document: chunking and embeddings that fit that document
  • A knowledge graph: entities and the relations between them
  • Several retrieval methods: vector, keyword + vector (hybrid), graph
  • A test set with scores, and a trace of what every query retrieved

Steps to attach it

  1. Upload the documents your agent must answer from.
  2. Write 20–50 real user questions, including ones that need two documents at once.
  3. Run them with plain vector retrieval and record the scores.
  4. Turn on the graph (or hybrid) for the questions that failed, and compare the scores.
  5. Keep the method that wins per use case; re-process documents without rebuilding the agent.
  6. Expose it to your agent through one tool: API, SDK, MCP or an n8n node.

💡 The trick: the questions that break plain RAG are relationship questions ("who approved it, and under which contract?"). Put at least five of them in your test set from day one. If they pass, the easy ones will too.

🧠 With GuidedMind.ai: this whole block is no-code. Upload documents, tune processing per document, switch between vector, hybrid and knowledge-graph retrieval, and test it with scores and traces before your agent depends on it. Your agent connects through the API, SDK, MCP or the n8n node, so you build the logic, not the plumbing. Try it free →

Go deeper

Chunk for the question, not the document. Small chunks (a few hundred tokens) win on precise lookups like a price or a policy line; section-sized chunks win on "explain" questions. That's why processing settings should be per document, not global.

Hybrid search catches what embeddings blur. Names, product codes, SKUs and error IDs are exactly where vector search is weakest. Keyword search (BM25) catches them, vector search catches paraphrases, and rank fusion combines both.

When the graph pays off. Multi-hop questions ("which clients are on plans that include X?") and global questions that summarize across many documents are where a knowledge graph beats similar-looking paragraphs. The cost moves to indexing time, when entities and relations are extracted, so plan how updates re-process documents.

Order and amount matter in the prompt. Models use facts at the start and end of a long context more reliably than facts in the middle. Send fewer, better facts, best first.

Grade before you trust. If retrieval scores low, don't answer anyway: rewrite the query and retry, or say you don't know. Corrective and self-reflective RAG formalize exactly this step.

📚 Read more: the research behind this block

2 · Decision layer: stop writing a condition for every scenario

Attach · business logic

Every agent has business logic: is this a booking, a refund, a complaint, a sales lead? Most teams start with an if-statement per scenario, then a switch node, then a giant prompt. Each works until the number of scenarios grows, and none of them tells you how sure the decision was. That missing number is what decides when a human should step in.

A message goes to a decision model with a list of options; it returns probability bars and a confidence above a 0.80 threshold, which routes automatically

A decision model returns a choice and how sure it is. The threshold line is where you hand off to a human.

OptionGood atBreaks when
If/else conditionsTwo or three fixed casesCustomers phrase things you didn't predict
Workflow switch nodesVisible routing in n8n or MakeEvery new scenario adds a branch and a keyword list
Decision tables (DMN)Auditable, rule-based domainsInputs are messy free text
LLM classifier promptFlexible, understands free textYou need to know how sure it was; outputs can drift
Decision model (e.g. Jev)One choice + probabilities + confidence, no textYou skip calibrating the threshold on real data

What's inside

  • A written list of options, one line of criteria each
  • A decision call that returns a choice, the probabilities and a confidence
  • A threshold that sends low-confidence or upset customers to a person
  • Thin branches: each one does one job

Steps to attach it

  1. List every scenario as an option with a one-line description.
  2. Add one yes/no question for "upset or asking for a person".
  3. Route on the chosen option; route anything under your threshold to a human.
  4. Run a few hundred past messages through it and check that 0.80 really means 80%.

How Jev works: Jev by TypeSafe AI is a decision model rather than a text model. You send the state (the customer message plus any facts) and your questions in one request. A choice question picks one option and returns the full probability distribution and a confidence; a score question returns a number; a yes/no question returns a probability. Several questions can go in one call and are evaluated in parallel. Jev is in early access, so treat its speed and price figures as TypeSafe's own and measure on your data.

💡 The trick: feed the decision the facts from your context layer, not just the message. "Customer on the Pro plan, contract renews Friday" turns an ambiguous message into an obvious one.

Go deeper

Calibrated means 0.8 is right about 80% of the time. Check it with a reliability table: bucket a few hundred past decisions by confidence (0.5–0.6, 0.6–0.7, …) and compare each bucket's confidence with its real accuracy. If 0.8 is only right 65% of the time, raise the threshold.

Pick the threshold with money, not taste. The threshold trades coverage (how much runs automatically) against risk (how often an automatic action is wrong). Set it where the cost of a wrong automatic action meets the cost of a person's minute. This is selective prediction: answer when sure, abstain otherwise.

Design the options like a form, not a vibe. Options should not overlap, each needs a one-line description, and there should be an "other" option. Bare labels without descriptions measurably lower confidence.

A text model's "I'm 90% sure" is not a probability. Confidence an LLM writes in its answer tends to be overconfident. A decision model that returns a probability distribution gives you a number you can calibrate.

📚 Read more: the research behind this block

3 · Tools: plug in MCP, don't wrap every API

Attach · actions

Writing a custom wrapper for every service is the slowest way to give an agent hands. Most services now ship an MCP server, and one MCP connection gives your agent a whole set of well-described tools. What's left for you is the part that matters: fewer, clearer tools.

What's inside

  • MCP servers for the services you use (chat, drive, calendar, databases)
  • Names in verb_noun form: search_knowledge, create_booking
  • Strict input schemas with required fields
  • Errors that say how to recover, not just "failed"

Steps to attach it

  1. List the actions the agent truly needs. Usually fewer than ten.
  2. Connect MCP servers for those services first.
  3. Wrap only what's missing, with a strict schema.
  4. Give knowledge lookups one tool that talks to your context layer.

💡 The trick: if two tools could answer the same request, the model will guess between them. Merge them or make their descriptions mutually exclusive.

Go deeper

Tool descriptions are prompts. The model decides from the name, the description and the schema alone. Say when to use the tool and when not to, and what it returns.

Return less, structured. A tool that returns 5,000 tokens of raw JSON burns context and buries the answer. Return the fields the agent needs, paginate the rest.

Make writes idempotent. Pass an idempotency key with every write, so a retry after a timeout can't book or charge twice.

Scope tools per task. Fewer visible tools means better tool choice. That's why each branch in the boilerplate sees only its own tools.

📚 Read more: the research behind this block

4 · Memory: three layers, not one giant chat history

Attach · continuity

Three memory layers: short-term session, long-term notes, and an entity graph of people, projects and clients

Dumping the whole history into every prompt is slow, expensive, and still forgets who's who.

What's inside

  • Short-term: the current session's recent turns
  • Long-term: summaries and facts worth keeping
  • An entity graph: people, projects, clients and how they relate

Steps to attach it

  1. Cap short-term memory at the last N turns.
  2. After each session, save a summary to long-term memory.
  3. Extract entities and relations into the graph as they appear.
  4. Retrieve from long-term and graph memory only when the question needs it.

🧠 With GuidedMind.ai: the GuidedMind n8n node includes short- and long-term agent memory next to the knowledge base, so memory and documents live in the same context layer.

Go deeper

Decide what's worth remembering. Store facts that change future behavior (preferences, commitments, relationships), not every message.

Update, don't overwrite. When a fact changes ("Anna moved to Beta Corp"), mark the old one invalid with a date instead of deleting it. Temporal knowledge graphs do exactly this, so the agent can answer both "where is she now?" and "where was she in March?"

Retrieve by relevance, recency and importance. A good memory lookup blends all three, not just similarity.

Memory is part of your eval set. Add questions that only work if something from a past session was remembered.

📚 Read more: the research behind this block

5 · Harness: the frame everything runs inside

Attach from day one · safety net

The harness is the runtime around your agent: it calls the model, runs the tools, retries what failed and stops what's running away. Without one, a single bad loop can burn a month of budget overnight, and a flaky API turns into a confused answer. You don't need to build it yourself. Agent SDKs (Claude Agent SDK, OpenAI Agents SDK), frameworks like LangGraph, and workflow tools like n8n all give you most of a harness. Your job is to set its limits on purpose.

The agent sits inside a harness frame with dials for retries, timeouts, sandbox, cost budget, rate limits and a kill switch

The harness sets the limits. The agent works inside them.

What's inside

  • Retries with backoff, and timeouts per call
  • A sandbox for any code the agent runs
  • A hard token and cost budget per conversation
  • Rate limits, and a kill switch you can flip
  • Idempotent tools, so a retry never books twice

Steps to attach it

  1. Pick a runtime (an agent SDK, LangGraph or n8n) instead of writing your own loop.
  2. Set a per-conversation cost ceiling before the first real user.
  3. Add retries only to calls that are safe to repeat.
  4. Wire an alert and a kill switch to the budget.

💡 The trick: set the cost ceiling per conversation, not per month. A monthly cap tells you about a runaway loop after it has already eaten the budget.

Go deeper

The harness owns the agent–computer interface. How tools present results, how errors read and how files are shown changes success rates as much as the model does. The SWE-agent research made this concrete for coding agents.

Manage the context window. Long runs need compaction: summarize old steps, truncate large tool outputs, keep the task and the latest state at the top.

Keep secrets out of prompts. Credentials live in the harness and are injected into tool calls, never into text the model can repeat.

Make runs replayable. Store inputs, tool results and model outputs per step, so you can re-run a failed conversation exactly when you debug.

📚 Read more: the research behind this block

6 · Loops: write down what "done" means

Attach · control

An agent loop is plan, act, check, repeat. Most loops that spiral never had a definition of done. Write one per task, and make the agent check it after every step.

What's inside

  • A maximum number of steps
  • A success check after each step
  • A written "done" definition per task
  • A no-progress detector: same action twice → stop

Steps to attach it

  1. Write the done definition in one sentence.
  2. Add a self-check step that compares the result to it.
  3. Stop and escalate on no progress, not just on the step limit.

Go deeper

Reason, act, observe, repeat. Interleaving reasoning with actions (ReAct) is the default loop for tool-using agents.

Reflection works when there's a real signal. Retrying after a failed test, a tool error or a failed eval check improves results. Asking the model to critique itself with no external signal often doesn't, and can make answers worse.

So make "done" checkable. A done definition the code can verify (a field is filled, a test passes, a booking ID exists) is worth more than "looks complete".

📚 Read more: the research behind this block

7 · Orchestration: climb the ladder only when you need to

Attach · structure

A ladder of four steps from one agent with tools up to a multi-agent supervisor

Each step up adds coordination cost. Most business agents never need the top step.

Start with one agent and a few tools. When a task has fixed steps, move it into a workflow. Add specialist agents only when tasks genuinely split and don't share context. Your decision layer (block 2) is often all the "orchestration" a business agent needs.

Go deeper

Workflows first, agents where you need judgment. Prompt chaining, routing, parallel calls, orchestrator–workers and evaluator–optimizer cover most business cases before you need autonomous multi-agent systems.

Multi-agent systems fail like organizations. Research on seven multi-agent frameworks found most failures come from unclear specifications, agents misaligned with each other, and weak verification, not from the individual models. Add agents only when tasks truly split and don't share context.

📚 Read more: the research behind this block

8 · Guardrails and approvals: scope every tool, approve what you can't undo

Attach · blast radius

What's inside

  • Per-tool scopes: only what that tool needs
  • Read tools separate from write tools
  • Allowlists for domains, recipients, tables
  • PII redaction before logging
  • Human approval for payments, sends and deletes

Steps to attach it

  1. Mark every tool as reversible or not.
  2. Put an approval step in front of every irreversible one.
  3. Route low-confidence decisions (block 2) to the same human queue.

Go deeper

Retrieved text is data, not instructions. Web pages, emails and documents can carry hidden instructions (indirect prompt injection). Never let retrieved content grant permissions or trigger writes on its own.

Put the gates at the tools. Separate read and write tools, allowlist destinations, and require approval for anything you can't undo. Output filters help, but the tool boundary is where damage happens.

Test attacks like you test features. Agent security benchmarks give you ready-made injection scenarios to run against your agent.

📚 Read more: the research behind this block

9 · Evals: a score, not a guess

Attach before shipping · proof

"It worked when I tried it" is how most agents ship. Evals replace that with a number you can compare before and after every change: new prompt, new model, new documents, new retrieval method.

What's inside

  • A golden set built from real user questions
  • Retrieval checks: did it fetch the right facts?
  • Outcome checks: did the task actually succeed?
  • Multi-document questions on purpose
  • Exact checks where possible, an LLM judge where not

Steps to attach it

  1. Collect questions from logs or the people who'll use it.
  2. Score retrieval separately from the final answer; they fail for different reasons.
  3. Re-run the set on every change and keep the history.

🧠 With GuidedMind.ai: retrieval testing is built in. Run your questions against the knowledge base, see similarity scores and traces, and compare retrieval methods before your agent depends on them. That's the half of evals most teams skip.

Go deeper

Evaluate three layers separately. Retrieval (did it fetch the right facts: context precision and recall), the decision (right branch, and is the confidence calibrated), and the end-to-end outcome (did the task succeed).

Test consistency, not one lucky run. Run the same conversation several times. An agent that succeeds once in three tries isn't reliable; τ-bench's pass^k metric measures exactly this.

LLM judges have biases. They favor longer answers and certain positions. Check a judge against a small set of human labels before trusting its scores.

📚 Read more: the research behind this block

10 · Observability: trace every answer back to its facts

Attach · debugging

A trace strip: context used, decision, tool calls, answer, each with cost and time

When an answer is wrong, you need to know which step went wrong: the facts it retrieved, the decision it made, or the tool it called. Trace all three, with time and cost per step. For the full run, tools like Langfuse, LangSmith, Arize Phoenix or plain OpenTelemetry work well. For the knowledge part, GuidedMind.ai traces which chunks and graph nodes every query used.

Go deeper

Capture per step: redacted inputs and outputs, model and version, tokens, latency, cost, tool arguments and results, retrieved chunk or node IDs, and the decision with its confidence.

Link traces to evals. Every failed eval case should open the trace that produced it. That's how you tell a retrieval miss from a wrong decision from a broken tool.

Use the standard. OpenTelemetry has semantic conventions for generative AI spans, so traces move between tools instead of being locked into one.

📚 Read more: the research behind this block

The one-page checklist

#BlockMinimum before real usersWhere
1Context layerTest set with multi-document questions, scored; retrieval method chosen per use caseGuidedMind.ai
2Decision layerOptions written down; confidence threshold calibrated on past messagesJev or similar
3ToolsMCP where possible; strict schemas; no overlapping toolsMCP servers
4MemoryShort-term capped; summaries saved; entities extractedGuidedMind.ai / others
5HarnessPer-conversation cost ceiling; retries only on safe calls; kill switchAgent SDK · LangGraph · n8n
6LoopsA written done definition and a no-progress stopYour runtime
7OrchestrationOne agent + workflows unless tasks truly splitn8n · your runtime
8GuardrailsEvery irreversible tool behind a human approvalYour runtime
9EvalsRetrieval and outcome scored separately; re-run on every changeGuidedMind.ai + eval tools
10ObservabilityEvery answer traceable to its facts, decision and tool callsGuidedMind.ai + tracing tools

Each row maps to one file in the boilerplate repo, so you can start from working code instead of a blank page.

Start with the block everything else depends on

Build the context layer without the data-layer rabbit hole: upload your documents, go from vector to hybrid to knowledge graph, and see a score before your agent ships.

Try GuidedMind.ai free → Get the boilerplate on GitHub

GuidedMind.ai is our product. Jev is made by TypeSafe AI and is in early access; we're not affiliated. The salon examples are illustrative, and the example numbers in the diagrams are for explanation, not benchmarks.