The short answer
An MCP tool result can carry two copies of the same answer: structuredContent, a JSON object for programs, and a text block for the model. The MCP specification says a tool with structured content SHOULD also put the serialized JSON in a text block, for backward compatibility (MCP specification 2025-06-18, Tools). We now break that SHOULD on purpose. A model that reads the text block was paying for every brace, every null field and every base64 picture.
| Across 113 tools | Before | After |
|---|---|---|
| Text sent back | 735,264 chars, about 183,816 tokens | 189,809 chars, about 47,452 tokens |
| Text blocks that were JSON | 109 of 116 | 0 |
| Empty-call refusals that list the parameters | 0 of 39 | 38 of 38 |
Agent got pick_ad_type right on the first try | 1 of 3 | 3 of 3 |
Token counts are characters divided by 4, a rough rule, not a tokenizer count.
What changed in the text block
- The text is Markdown written from
structuredContent, not a copy of it. - Null and empty fields are left out, and a field that is the same in every row is said once.
- Short rows become a table.
- Base64, SVG and HTML blobs are named by size and place, for example "HTML page, 42 KB, in structuredContent.html", instead of copied.
- Six tools with answers that read badly in a generic layout got their own short text writer.
No tool name, parameter, schema or description changed, and structuredContent stayed the same. tools/list on all 40 hosts matched the snapshot taken before the change.
Where the characters went
| Tool | Text before | Text after |
|---|---|---|
draft_tiktok_slides | 481,286 | 3,694 |
chart | 44,142 | 1,544 |
launch_score, Mailcheck | 10,192 | 4,110 |
get_tax_parameters, Tax | 8,635 | 2,339 |
find_company_jobs, Company Jobs | 5,219 | 1,681 |
list_recalls | 5,231 | 5,098 |
Two tools did most of the work. The slides answer carried PNG pictures as base64 text, 481,286 characters for one call. The chart answer carried a whole HTML page. Most tools changed little. Of the 104 tools that answered without an error both times, the median text fell by about 5%, and only 6 fell by half or more. So the saving comes from a few answers that carried blobs, not from trimming every answer. The recall tools barely moved, and that is correct: their text is the recall summaries, which is the answer itself.
Why less text helps a model, not only a bill
Long input does not just cost more. Models use facts in the middle of a long context worse than facts at the start or the end (Liu and others, Lost in the Middle, 2023). Chroma tested 18 models and found that performance falls as input grows, even on simple tasks (Hong, Troynikov and Huber, Context Rot, 2025). Anthropic's own guide to tool design asks for token-efficient answers and notes that Claude Code limits a tool response to 25,000 tokens by default (Anthropic, Writing effective tools for AI agents, 2025). Our old slides answer was about 120,000 tokens of text.
Who gains, and who does not
This is the part we did not expect. Claude Code passes only structuredContent to the model when a result has it. In our agent check, a Haiku agent in Claude Code got the same 10,192 characters for launch_score before and after the change. So the smaller text saves Claude Code users nothing. ChatGPT and other clients that read the text block get the full drop.
The biggest blobs are still in structuredContent, because their fields are declared and required in the published output schema. The slides tool still sends about 481 KB of base64 there, and Claude Code reads it in full. Moving those blobs to a link waits for a schema revision.
The change that helped every client: refusals that say what to send
A refusal has no structuredContent, so its text reaches every client, Claude Code included. Before, an empty call got invalid_input and nothing else. Now it lists the parameters: invalid_input: domains is required. Parameters: domains (array, required). An unknown tool name now lists the real tools, and a server failure ends with a line on how to report it.
The clearest result was pick_ad_type. Its schema allows a call with only a product, but the tool refused it. In 3 trials before the change, 2 agents sent only the product, were refused, and stopped to ask the user. After it, the tool answers with the two plays for a cold audience and a note on what to add, and 3 of 3 agents finished on the first try. That is a small sample, so read it as a direction, not a rate.
What others found
A 2026 study of 856 tools on 103 MCP servers found at least one description smell in 97.1% of them, and an unclear purpose in 56% (Hasan and others, 2026). Rewriting the descriptions raised task success by a median of 5.85 points but took 67.46% more steps. Another 2026 study found that running code inside the MCP server cuts tokens and latency but opens a much wider attack surface (Felendler and others, 2026). On X, Dhravya Shah described the same lesson at supermemory: a document fetch had grown to 23 fields, including internal ids that were never meant to be public, and endpoints called by agents now answer in Markdown (Shah, 2026). That post was the starting idea for our change.
How to measure your own server
- Pick one real input per tool and keep it fixed.
- Call every tool once and add up the length of the text blocks. Count the ones that start with
{. - Call every tool once with empty arguments and read the refusals: does each one say what to send?
- Change the text writer, run the same calls, and compare. Then diff tools/list before and after, so the change never touches your listing.
Try it: every Agent Tools server answers this way. Browse the install guides and add one to your client.
Sources
- MCP specification 2025-06-18, Tools: a tool with structuredContent SHOULD also return the serialized JSON as text.
- Ken Aizawa, Anthropic. Writing effective tools for AI agents (2025): return token-efficient answers; Claude Code caps a tool response at 25,000 tokens by default.
- Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni, Liang. Lost in the Middle (2023): facts in the middle of a long context are used worst.
- Hong, Troynikov, Huber. Context Rot (Chroma, 2025): performance falls as input length grows, across 18 models.
- Hasan, Li, Rajbahadur, Adams, Hassan. MCP Tool Descriptions Are Smelly! (2026): 97.1% of 856 tools have a description smell.
- Felendler, Gandhi, Habler, Elovici, Shabtai. From Tool Orchestration to Code Execution (2026): code-execution MCP saves tokens but widens the attack surface.
- Dhravya Shah on X. Designing an Agent-first API (2026): a 23-field document answer full of internal ids, and Markdown answers for agents.
- Our own audit: 113 Agent Tools tools, same inputs, live hosts before and a local build after, 2026-10-07; agent check with Claude Haiku 4.5 in Claude Code, 3 trials per tool and side.