Blog
Why frontier models handle heavy tool and skill loads, and smaller models struggle
Agents get worse as you add tools, skills and history. The benchmarks, the causes, and what it means when you pick an app for your Claude or ChatGPT plan.
By BringYourSubs6 min read
Agents get worse as you hand them more tools, more skills and longer histories, and the drop is steepest for smaller and older models. Frontier models degrade too, but on benchmarks built on real tool servers they start higher and fall less, so an app that loads dozens of tools or skills needs a frontier model behind it. That is a large part of why running apps on your own Claude or ChatGPT plan matters: the plan gives you frontier models at a flat price.
Below is the evidence, why it happens, where the claim needs care, and what to look for when you choose an app.
What happens to an agent when you give it more tools?
It gets slower to start and worse at choosing. Three measurements show the effect from different angles.
Tool definitions eat the context before work begins. Anthropic counted a five-server setup (GitHub, Slack, Sentry, Grafana, Splunk) at 58 tools and about 55,000 tokens "before the conversation even starts", and says it has seen tool definitions reach 134,000 tokens internally. It also names the failure modes: "The most common failures are wrong tool selection and incorrect parameters, especially when tools have similar names."
Accuracy falls as the catalogue grows. IBM Research's LongFuncEval grew the tool catalogue from 8,000 to 120,000 tokens and measured "a performance drop of 7% to 85% as the number of tools increases", plus 7% to 91% degradation as tool responses got longer and 13% to 40% as conversations got longer.
Showing fewer tools fixes much of it. The RAG-MCP paper by Tiantian Gan and Qiyao Sun retrieved only the relevant tools before calling the model and raised tool-selection accuracy from 13.62% to 43.13%. Anthropic's own Tool Search Tool, which loads tools on demand, took Opus 4 from 49% to 74% on its internal MCP evaluation with a large tool library, and Opus 4.5 from 79.5% to 88.1%.
Look at those last two numbers together. The older model gained 25 points from a cleaner context; the newer one gained under 9 because it was already coping with the load. That is the pattern this post is about.
Do frontier models really do better on agent benchmarks?
On benchmarks that use live MCP servers rather than mocked functions, yes, though nobody scores close to perfect.
| Benchmark | What it tests | Strongest result reported | Weaker results reported |
|---|---|---|---|
| MCP-Universe (Salesforce, Aug 2025) | 11 real MCP servers across 6 domains, long multi-step tasks | GPT-5 at 43.72% | Grok-4 33.33%, Claude-4.0-Sonnet 29.44% |
| MCPMark (Sep 2025) | 127 create/read/update/delete tasks on Notion, GitHub, Postgres and more | gpt-5-medium at 52.56% pass@1, 33.86% pass^4 | Authors say most models stay under 30% pass@1; open models such as Qwen-3-Coder and K2 around 20% |
| MCP-Atlas (Scale AI, 2026) | 1,000 tasks over 36 servers and 220 tools, 6 to 37 tools exposed per task with only 2 to 8 relevant | Top tier 78.2% to 82.2%, including Claude Opus 4.7 at 79.1% | Long tail down to 40.2%; o3 Pro at 44.5% |
| Anthropic internal MCP eval, large tool library | Tool selection and use at scale | Opus 4.5 at 88.1% with Tool Search | Opus 4 at 49% with all tools loaded |
Two things stand out. The spread between the top and bottom of each table is wide, often 30 to 40 points. And MCP-Atlas found what it calls "a clear three-tier performance structure": a top cluster of three models, a mid-tier of eight between 67.6% and 76.8%, and a long tail of nine. MCP-Atlas built its tasks on purpose so the agent has to find the right tools "among semantically plausible distractors", which is exactly the condition an app with a big skill or tool library creates.
Why do agents get worse as the load grows?
Several separate mechanisms stack up.
The attention budget gets spread thin. Anthropic's engineering team describes context as "a finite resource with diminishing marginal returns". Transformers relate every token to every other token, "n² pairwise relationships for n tokens", and models learn their attention patterns mostly from shorter sequences. The result, in Anthropic's words, is "a performance gradient rather than a hard cliff".
Length hurts even when retrieval is perfect. Yufeng Du, Hao Peng and co-authors ran controlled experiments on five models and found performance still fell 13.9% to 85% as input grew, well inside the advertised context windows. It happened even when the filler was whitespace, and even when the irrelevant tokens were masked out.
Models do not use their context evenly. Chroma's Context Rot report by Kelly Hong, Anton Troynikov and Jeff Huber tested 18 models with the task held constant and only length varied. In Chroma's own summary on X: "models do not use their context uniformly." Real agent work is messier than their controlled tasks, which the authors say likely makes the effect larger in practice.
Look-alike tools invite wrong picks. Anthropic's example is notification-send-user versus notification-send-channel. Every extra server adds near-duplicates.
Long horizons pile on more tokens. The MCP-Universe authors note that "the number of input tokens increases rapidly with the number of interaction steps", and that agents often lack familiarity with the exact usage of unfamiliar MCP servers.
Skills follow the same logic. Anthropic designed Agent Skills around progressive disclosure: only each skill's name and description sit in the system prompt, and the full instructions load when a task needs them. That design exists because loading every skill body upfront would hit the same wall.
What do people running local and smaller models report?
The forums say the same thing as the papers, in plainer words.
In an r/LocalLLaMA thread titled "Anyone else have small models just forget MCP tools exist?", the poster wrote: "Most of the time, Qwen doesn't even seem to know that the MCP tools are there." In another thread on struggles with local tool calling, one commenter exploring "tool calling with 100s of tools" said "that just fails aggressively". The same thread names Qwen 3 thinking models and GPT-OSS 120b as the local models that held up best, and notes that results drop when a non-reasoning model fires a tool call straight after the user's request.
Academic work on small models agrees. Less is More by Varatheepan Paramanayakam and colleagues found that "selectively reducing the number of tools available to LLMs significantly improves their function-calling performance" on edge hardware, while cutting execution time by up to 70% and power by up to 40%.
Where does this claim need nuance?
"Use a frontier model" is a good default, not a law.
- Strong reasoning is not the same as strong tool use. In MCP-Atlas, o3 Pro, a leading model on maths and coding, finished near the bottom at 44.5% because it made no tool calls on 40% of the tasks it failed.
- Open models are closing in. GLM-5.1 scored 75.6% on MCP-Atlas, which the authors say enters a band "previously exclusive to proprietary models".
- Many failures are not tool mistakes. MCP-Atlas diagnosed 63.3% of failures as cognitive: misreading the task, poor synthesis, stopping early.
- The harness is not a cure. MCP-Universe found that "enterprise-level agents like Cursor cannot achieve better performance than standard ReAct frameworks."
- Benchmarks age fast. MCP-Universe and MCPMark tested 2025 models. Today's models score higher, so read the gaps, not the absolute numbers.
- Frontier models still rot. Anthropic says context degradation "emerges across all models". A frontier model in a bloated context can lose to a smaller model in a clean one.
What does this mean when you pick an app?
- Check which model does the heavy lifting. If an app loads many tools or skills, it should run a frontier model. Claude Pro and Max include Opus; ChatGPT Plus includes GPT-6.1 Sol in Codex. Running the app on your Claude plan or your ChatGPT plan is the cheapest way to get those models for agent work.
- Prefer apps that load tools and skills on demand. Tool search, progressive disclosure and retrieval all keep the context small. Ask the developer, or check the docs.
- Prefer several focused agents over one overloaded one. Apps such as Conductor, Munder Difflin and T3 Code run separate Claude Code or Codex sessions, each with its own context.
- Keep sessions short. Start fresh when a task changes rather than letting one conversation grow for hours.
- Use small or local models where the job is narrow. A single-purpose tool with four distinct functions does not need the biggest model.
For the cost side of this decision, read why using your own Claude or ChatGPT subscription matters.
Questions people ask
- Why does my agent get worse when I add more MCP servers?
- Every tool definition is loaded into the model's context, so more servers mean more tokens and more look-alike options. Anthropic counted 58 tools across five common MCP servers at about 55,000 tokens, and says the most common failures are wrong tool selection and incorrect parameters. IBM's LongFuncEval measured drops of 7% to 85% as the tool catalogue grew.
- What is the best model for tool calling?
- On real MCP-server benchmarks the leaders are frontier models. In Scale AI's MCP-Atlas the top three scored 78.2% to 82.2%, and the best open model, GLM-5.1, reached 75.6%. Rankings change with every release, so check a current tool-use benchmark rather than a general leaderboard.
- Do frontier models also suffer from context rot?
- Yes. Chroma tested 18 models and found performance degrades as input length grows, and Anthropic says the effect appears across all models, with some degrading more gently than others. Frontier models start higher and hold up better, but they still benefit from smaller, cleaner context.
- Can small or local models run agents?
- Yes, for narrow jobs with a handful of distinct tools. Research such as Less is More shows small models do much better when they are shown fewer tools. They struggle when an app loads dozens of tools or long skill files at once.
Sources
- Introducing advanced tool use on the Claude Developer Platform, Anthropic, 2025-11-24
- Effective context engineering for AI agents, Anthropic, 2025-09-29
- Equipping agents for the real world with Agent Skills, Anthropic, 2025-10-16
- Context Rot: How Increasing Input Tokens Impacts LLM Performance, Chroma (Hong, Troynikov, Huber), 2025-07-14
- Chroma announcement of the Context Rot report, Chroma on X, 2025-07-14
- MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers, Salesforce AI Research (arXiv), 2025-08-20
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use, arXiv, 2025-09-28
- Introducing MCPMark, MCPMark, 2025-08-26
- MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers, Scale AI (arXiv), 2026-10-03
- LongFuncEval: Measuring the effectiveness of long context models for function calling, IBM Research (arXiv), 2025-04-30
- RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation, arXiv, 2025-05-06
- Context Length Alone Hurts LLM Performance Despite Perfect Retrieval, arXiv, 2025-10-06
- Less is More: Optimizing Function Calling for LLM Execution on Edge Devices, arXiv (DATE 2025), 2024-11-23
- Anyone else have small models just forget MCP tools exist?, Reddit r/LocalLLaMA, 2025-09-15
- What are your struggles with tool-calling and local models?, Reddit r/LocalLLaMA, 2025-09-01
- Plans & Pricing, Anthropic, 2026-10-03
- Pricing (ChatGPT Work and Codex), OpenAI, 2026-10-03