The right skill on the right turn, or none at all

Skills6 min read

You wrote skills so your agent follows your process: how your team debugs, splits a plan into tickets, reviews a diff. But a skill only helps if it loads on the right turn and stays out when you are just chatting. Since October 8, 2026, Context Mode makes that choice by asking Cloudflare Clef one question over your whole skill catalog. On a small offline sample, correct picks went from 4 of 10 to 10 of 10, and the slowest choices got much faster. This post shows how it works, what we measured, and what we have not measured yet.

The problem: skills load at the wrong time

A skill is a block of instructions. On each turn, the choice can go wrong in three ways:

Storing skills is easy. The hard part is the choice on each turn: this skill, a different one, or nothing.

One catalog, on every machine

Claude Code and Codex load skills from SKILL.md files on your local disk. Context Mode keeps them in the gateway instead. You import a GitHub repo once in the console, and every session that goes through Context Mode sees the same catalog, on any machine. Nothing is written to your disk.

Each enabled skill adds one line to a list the model sees on every request. In our benchmark accounts that list was about 760 tokens a request, read from cache. A skill that loads adds its whole body, at most 262,144 bytes per skill. That cost is why the choice matters.

How Context Mode picks a skill

It checks three paths, in this order.

  1. You named the skill. If your client sends a skill name that is in your catalog, that skill loads. No score or model gets a vote. If you name a skill inside a sentence ("run batch-grill-me on the todo app"), a separate check fires it when you ask for it with a target, and stays quiet when you only talk about it ("can we disable it").
  2. You did not name one. Context Mode asks Cloudflare Clef, a 27B judge model, one question: which of your enabled skills fits this message, or NONE? Each option carries the skill's description. Clef gives a probability for every option, and a skill fires only if Clef picks it with p >= 0.5.
  3. Large catalogs. With more than 24 skills, a fast embedding search first makes a shortlist of at most 8, and Clef picks from that.

NONE has its own written rule, so "no skill" is a real answer and not a fallback:

No listed capability: the message asks for information or an explanation,
gives a status update, acknowledges, talks ABOUT a capability, or orders
something no listed capability does.

At most one skill loads per turn. When one does, a line in the reply tells you which:

context-mode · skill to-tickets auto-loaded

The proof: before and after

Before October 8, the judge was llama-3.3-70B, and it only saw the skills an embedding search found first. We tested both judges offline on recorded turns and found two problems with the old one.

It never saw some right answers. On a labeled set of 41 made-up cases, 3 of the old route's 4 misses came from the shortlist, not the judge. The embedding search found no candidate for a Turkish UI request, a Turkish plan-review request and an English "break this plan into tickets". If the right skill is not on the list, no judge can pick it. Clef over the whole catalog got all 3.

Its picks were often wrong. We took the 12 distinct texts where the two judges disagreed, left out 2 that the routing rule does not decide, and judged the other 10 by hand. All 3 of the old judge's fires in that sample were wrong, and it missed clear UI-design and bug-fix requests.

llama-3.3-70B (before)4/10
Clef, shortlist, p >= 0.59/10
Clef, whole catalog, p >= 0.510/10
Correct choices on the 10 texts where the judges disagreed. Offline eval, small sample, judged by one person.
JudgeShortlistWhole catalog
llama-3.3-70B (before)4/104/10
Clef 27B, top answer7/106/10
Clef 27B, p >= 0.59/1010/10

On the 41-case labeled set, the old route got 37/41 and Clef over the whole catalog got 39/41. The confidence ranges overlap (0.78 to 0.96 and 0.84 to 0.99), so that set does not separate the two. Read it as a check that nothing got worse, not as proof of a gain.

Why "only when sure" matters

The 0.5 floor is what stops skills from loading on chatter. Without it, Clef's top answer fired on one-word test prompts at p 0.3 to 0.4. In that form it fired on 21 of 121 turns, and per turn it was right less often than the old judge (10/24 vs 16/24). With the floor, it fired on 9.

We also tested the smaller Clef-flash (9B). It is cheaper, but it picked a wrong skill on UI requests, so we did not ship it for this choice.

It does not slow your turn

The chosen skill goes into this turn's prompt, so the choice sits on the request path. Three things keep it short:

llama-3.3-70B, p954,728 ms
Clef 27B, p95612 ms
p50 was 397 ms for the old judge and 281 ms for Clef. n 20 per judge.

From outside over REST (which adds about 300 to 550 ms of network), 9 of 121 old-judge calls took over 2,000 ms, and the longest took 42.9 s. Clef had 5. We have not measured the whole-catalog form inside the Worker; over REST it was p50 660 ms and p95 1,523 ms.

A timeout is no longer called a "no"

With the old judge, a call over 2,000 ms was stopped and logged as judge-none, the same label as a real "no skill". You could not tell "the model said no" from "the model never answered". Now each case has its own label:

judge-fire        a skill passed the floor and loaded
judge-none(clef)  Clef answered NONE, or its pick was under 0.5
judge-timeout     no answer within 2,000 ms
judge-error(...)  error, malformed answer, or the model was unavailable

A timeout or an error loads nothing, and it never counts as a "no".

What it costs you

Nothing extra. Context Mode runs the judge on its own Cloudflare account. Your Anthropic or OpenAI account pays only for the model call it was already making, plus the skill body when one loads.

For the record, Clef reads more input than the old judge (848 vs 683 tokens) but has no output charge. In the eval it cost 0.204 USD per 1,000 calls on the shortlist and 0.230 over an 11-skill catalog, against 0.214 for the old judge.

What we have not measured

How to try it

  1. Run npx @context-mode/cli, sign in, connect Claude Code or Codex, and restart the client. The quick start has each step.
  2. In the console, open Skills, paste a GitHub repo such as mattpocock/skills, click Preview import, then Add.
  3. Describe a task in plain words without naming a skill. The reply line tells you which skill loaded, or none.

Read more on the Skills page and in the Skills docs. The Free plan needs no card; see plans and limits for what each plan includes.