# The right skill on the right turn, or none at all

October 9, 2026 Skills 6 min read

You wrote skills so your agent follows your process: how your team debugs, splits a plan into tickets, reviews a diff. But a skill only helps if it loads on the right turn and stays out when you are just chatting. Since October 8, 2026, Context Mode makes that choice by asking Cloudflare Clef one question over your whole skill catalog. On a small offline sample, correct picks went from 4 of 10 to 10 of 10, and the slowest choices got much faster. This post shows how it works, what we measured, and what we have not measured yet.

## The problem: skills load at the wrong time

A skill is a block of instructions. On each turn, the choice can go wrong in three ways:

- **The wrong skill loads.** You pay for a large prompt block, and it pulls the agent off task.
- **No skill loads when one was needed.** Your process is not followed, so the skill might as well not exist.
- **A skill loads on chatter.** A greeting, a status update or a one-word reply drags a whole skill body into the prompt.

Storing skills is easy. The hard part is the choice on each turn: this skill, a different one, or nothing.

## One catalog, on every machine

Claude Code and Codex load skills from `SKILL.md` files on your local disk. Context Mode keeps them in the gateway instead. You import a GitHub repo once in the console, and every session that goes through Context Mode sees the same catalog, on any machine. Nothing is written to your disk.

Each enabled skill adds one line to a list the model sees on every request. In our benchmark accounts that list was about 760 tokens a request, read from cache. A skill that loads adds its whole body, at most 262,144 bytes per skill. That cost is why the choice matters.

## How Context Mode picks a skill

It checks three paths, in this order.

1. **You named the skill.** If your client sends a skill name that is in your catalog, that skill loads. No score or model gets a vote. If you name a skill inside a sentence ("run batch-grill-me on the todo app"), a separate check fires it when you ask for it with a target, and stays quiet when you only talk about it ("can we disable it").
2. **You did not name one.** Context Mode asks Cloudflare Clef, a 27B judge model, one question: which of your enabled skills fits this message, or `NONE`? Each option carries the skill's description. Clef gives a probability for every option, and a skill fires only if Clef picks it with p >= 0.5.
3. **Large catalogs.** With more than 24 skills, a fast embedding search first makes a shortlist of at most 8, and Clef picks from that.

`NONE` has its own written rule, so "no skill" is a real answer and not a fallback:

```
No listed capability: the message asks for information or an explanation,
gives a status update, acknowledges, talks ABOUT a capability, or orders
something no listed capability does.
```

At most one skill loads per turn. When one does, a line in the reply tells you which:

```
context-mode · skill to-tickets auto-loaded
```

## The proof: before and after

Before October 8, the judge was llama-3.3-70B, and it only saw the skills an embedding search found first. We tested both judges offline on recorded turns and found two problems with the old one.

**It never saw some right answers.** On a labeled set of 41 made-up cases, 3 of the old route's 4 misses came from the shortlist, not the judge. The embedding search found no candidate for a Turkish UI request, a Turkish plan-review request and an English "break this plan into tickets". If the right skill is not on the list, no judge can pick it. Clef over the whole catalog got all 3.

**Its picks were often wrong.** We took the 12 distinct texts where the two judges disagreed, left out 2 that the routing rule does not decide, and judged the other 10 by hand. All 3 of the old judge's fires in that sample were wrong, and it missed clear UI-design and bug-fix requests.

**llama-3.3-70B (before)** *4/10* **Clef, shortlist, p >= 0.5** *9/10* **Clef, whole catalog, p >= 0.5** *10/10*

Correct choices on the 10 texts where the judges disagreed. Offline eval, small sample, judged by one person.

| Judge | Shortlist | Whole catalog |
| --- | --- | --- |
| llama-3.3-70B (before) | 4/10 | 4/10 |
| Clef 27B, top answer | 7/10 | 6/10 |
| Clef 27B, p >= 0.5 | 9/10 | 10/10 |

On the 41-case labeled set, the old route got 37/41 and Clef over the whole catalog got 39/41. The confidence ranges overlap (0.78 to 0.96 and 0.84 to 0.99), so that set does not separate the two. Read it as a check that nothing got worse, not as proof of a gain.

## Why "only when sure" matters

The 0.5 floor is what stops skills from loading on chatter. Without it, Clef's top answer fired on one-word test prompts at p 0.3 to 0.4. In that form it fired on 21 of 121 turns, and per turn it was right less often than the old judge (10/24 vs 16/24). With the floor, it fired on 9.

We also tested the smaller Clef-flash (9B). It is cheaper, but it picked a wrong skill on UI requests, so we did not ship it for this choice.

## It does not slow your turn

The chosen skill goes into this turn's prompt, so the choice sits on the request path. Three things keep it short:

- **It runs in parallel.** The whole-catalog question needs no shortlist, so Clef starts at the same time as the embedding search.
- **A hard stop.** The Clef call stops at 2,000 ms. The whole skill-match step has a 4,000 ms limit.
- **A shorter slow tail.** Measured inside the Worker, 20 calls each, shortlist form:

**llama-3.3-70B, p95** *4,728 ms* **Clef 27B, p95** *612 ms*

p50 was 397 ms for the old judge and 281 ms for Clef. n 20 per judge.

From outside over REST (which adds about 300 to 550 ms of network), 9 of 121 old-judge calls took over 2,000 ms, and the longest took 42.9 s. Clef had 5. We have not measured the whole-catalog form inside the Worker; over REST it was p50 660 ms and p95 1,523 ms.

## A timeout is no longer called a "no"

With the old judge, a call over 2,000 ms was stopped and logged as `judge-none`, the same label as a real "no skill". You could not tell "the model said no" from "the model never answered". Now each case has its own label:

```
judge-fire        a skill passed the floor and loaded
judge-none(clef)  Clef answered NONE, or its pick was under 0.5
judge-timeout     no answer within 2,000 ms
judge-error(...)  error, malformed answer, or the model was unavailable
```

A timeout or an error loads nothing, and it never counts as a "no".

## What it costs you

Nothing extra. Context Mode runs the judge on its own Cloudflare account. Your Anthropic or OpenAI account pays only for the model call it was already making, plus the skill body when one loads.

For the record, Clef reads more input than the old judge (848 vs 683 tokens) but has no output charge. In the eval it cost 0.204 USD per 1,000 calls on the shortlist and 0.230 over an 11-skill catalog, against 0.214 for the old judge.

## What we have not measured

- **Whether skills give better answers.** This post is about choosing the skill, not about what the skill does.
- **Clef accuracy on live traffic.** The 9/10 and 10/10 come from an offline eval: one person judged 10 distinct texts after seeing the model answers. The 0.5 floor was chosen after looking at the data. The test catalog had 11 skills, not a real customer catalog. Treat the numbers as direction.
- **Talking about a skill.** On the labeled set, both judges fired on a question about why a skill fired. That case is not solved.

## How to try it

1. Run `npx @context-mode/cli`, sign in, connect Claude Code or Codex, and restart the client. The [quick start](https://context-mode.com/docs/quick-start) has each step.
2. In the console, open Skills, paste a GitHub repo such as `mattpocock/skills`, click **Preview import**, then **Add**.
3. Describe a task in plain words without naming a skill. The reply line tells you which skill loaded, or none.

Read more on the [Skills page](https://context-mode.com/skills) and in the [Skills docs](https://context-mode.com/docs/skills). The Free plan needs no card; see [plans and limits](https://context-mode.com/docs/plans-and-limits) for what each plan includes.

[Start free](https://context-mode.com/docs/quick-start)[All posts](https://context-mode.com/blog)
