direkt zum inhalt

29. September 2026kiunternehmen

Drei Skills für Claude Code geprüft: was ich einschalte und was aus bleibtThree Claude Code skills tested: what I switch on and what stays off

context-mode, VibeSec und stop-slop versprechen weniger Token, sichereren Code und bessere Texte. Ich habe alle drei gegen mein Setup geprüft. Jedes Werkzeug bekam eine andere Entscheidung.context-mode, VibeSec and stop-slop promise fewer tokens, safer code and better writing. I tested all three against my own setup. Each tool got a different decision.

Eine Balkenwaage aus Messing auf einem dunklen Tisch im warmen Licht. In einer Schale liegen zwei kleine Gewichte, ein größeres steht daneben auf dem Tisch.
Titelbild KI-generiert, als solches gekennzeichnet.

Heute kamen drei Links auf meinen Tisch. context-mode verspricht 98 Prozent weniger Kontext. VibeSec soll Claude Code dazu bringen, Web-Code wie ein Bug Hunter zu schreiben. stop-slop räumt KI-Muster aus Texten. Die Bitte dazu war klar: alle drei installieren und ins bestehende System einbauen, für bessere Ergebnisse und weniger Tokenverbrauch.

Am Ende bekam jedes der drei Werkzeuge eine andere Entscheidung. VibeSec sitzt jetzt fest im Ablauf. context-mode bleibt aus und lässt sich pro Sitzung zuschalten. stop-slop liegt installiert im Ordner und hängt an keiner Regel.

Dieser Beitrag gehört zur Reihe über mein Claude-Code-Setup. „Claude Code einrichten“ baut das Grundsystem, „Die Agent-Schleife“ beschreibt den Loop mit drei Runden, „Claude Code baut. Codex prüft“ ergänzt das Review. Du kannst ihn auch ohne die anderen Teile lesen.

der kern in vier sätzen

Ein Werkzeug, das Token sparen soll, muss im eigenen Setup gegen einen Lauf ohne es antreten. context-mode war bei mir 11 Prozent teurer, weil lean-ctx die Arbeit schon macht. VibeSec kostet nur dann Kontext, wenn Web-Code mit Sicherheitsbezug entsteht, und genau dort gehört es hin. Für stop-slop gab es keine Lücke.

context-mode im a/b+11 %teurer pro lauf, gleiche antwort
headroom-proxy0,7 %kompression nach lean-ctx
sockel pro sitzung17.000token, vorher 73.000

das system vor dem update

Mein Setup hat einen festen Ablauf von DEFINE bis SHIP. Das Hauptmodell plant und nimmt ab, günstige Instanzen bauen die einzelnen Schritte, Codex prüft als unabhängiger Reviewer. Jede Sitzung startet mit lean-ctx, einem lokalen Werkzeug, das Datei-Lesen, Suchen und Shell-Ausgaben komprimiert, bevor sie im Kontext landen.

Der feste Vorspann einer Sitzung lag bis zum 10.09. bei rund 73.000 Token. Nach einer Aufräumrunde an diesem Tag waren es 17.000. Seitdem gilt eine einfache Regel. Was in den Startkontext will, muss dort messbar mehr sparen, als es kostet.

Das Schaubild zeigt, wo die drei neuen Werkzeuge andocken. Tippe auf einen Knoten, dann siehst du seine Verbindungen.

Das System in vier Schichten: Sitzung mit lean-ctx und optional context-mode, Orchestrierung mit orchestrate und executor, Prüfung mit VibeSec, Security-Hook und Codex, Text mit humanizer und stop-slop.

  • sitzung: lean-ctx (Komprimiert Lesen, Suchen und Shell-Ausgaben in jeder Sitzung. Standard.); context-mode (Sandbox und Suchindex für große Ausgaben. Nur per claude-cm, nur für eine Sitzung.)
  • orchestrierung: orchestrate (Plan mit Prüfbefehl je Schritt, harte Grenze von drei Runden.); executor (Günstige Instanzen bauen die Schritte, das Hauptmodell nimmt ab.)
  • prüfung: VibeSec (Neu: Regeln für sicheren Web-Code, geladen beim Schreiben.); security-hook (Prüft jede Änderung nach dem Schreiben.); Codex (Unabhängiges Review, adversarial bei Sicherheitsthemen.)
  • text: humanizer (Pflicht für deutsche Texte mit Außenwirkung.); stop-slop (Neu installiert, bewusst in keine Regel eingebaut.)

Verbindungen: von lean-ctx zu orchestrate, von context-mode zu executor, von orchestrate zu executor, von executor zu VibeSec, von VibeSec zu security-hook, von security-hook zu Codex, von executor zu humanizer.

Tippe auf einen Knoten, um seine Verbindungen zu sehen.

context-mode: gemessen und für teuer befunden

context-mode ist ein Plugin mit eigenem MCP-Server. Es führt Code in einer Sandbox aus, legt große Ausgaben in einen lokalen Suchindex und gibt dem Modell nur die passenden Treffer zurück. Dazu hängt es Hooks an Read, Grep, WebFetch, Agent und jedes MCP-Tool. Die Idee ist gut. Rohe Ausgaben fressen in langen Sitzungen den größten Teil des Kontexts.

Installiert hatte ich es schon am 18.09. Am selben Tag lief der A/B-Test. Die Aufgabe war bewusst ausgabelastig. Claude sollte die drei meistgeänderten Dateien aus git log --stat eines Repos mit rund 168 KB Ausgabe finden. Vier Läufe ohne context-mode, zwei mit, jeder Arm per --settings fest eingestellt.

kosten je lauf in us-dollar, dieselbe aufgabe mit context-mode aus und an
KategorieWert
aus 11,28 $
aus 21,26 $
aus 31,16 $
aus 41,02 $
an 11,30 $
an 21,31 $
  • aus 11,28 $
  • aus 21,26 $
  • aus 31,16 $
  • aus 41,02 $
  • an 11,30 $
  • an 21,31 $

Ohne das Plugin kosteten die Läufe zwischen 1,02 und 1,28 Dollar. Mit Plugin lagen beide Läufe bei 1,30 und 1,31 Dollar, im Mittel rund 11 Prozent darüber. Alle sechs Antworten waren richtig.

Den Grund zeigt der Cache. Mit context-mode schrieb jede Sitzung rund 116.000 Token in den Cache, ohne waren es 94.000 bis 113.000. Das Plugin bringt elf Tools und einen Block mit Routing-Regeln in den Vorspann. Die Kompression, die es anbietet, hatte lean-ctx schon erledigt.

cache_write in tausend token: mit context-mode wächst der feste vorspann jeder sitzung
KategorieWert
aus, min94k
aus, max113k
an116k
  • aus, min94k
  • aus, max113k
  • an116k

Dazu kam eine Falle. Der Hook von context-mode ersetzt curl-Aufrufe ohne -s -o datei durch ein echo. Eine Live-Prüfung wie curl -sI auf einen Cache-Header liefert dann still nichts zurück, und du suchst den Fehler auf dem Server.

Einen Tag vorher hatte ich mit einem Kompressions-Proxy dasselbe Muster gesehen. Er komprimierte nach lean-ctx noch 0,7 Prozent und machte den Lauf 2,5 Prozent teurer. Zwei Werkzeuge für dieselbe Aufgabe sparen nicht doppelt. Das zweite zahlt nur noch Eintritt.

die lösung: eine weiche statt eines schalters

Global einschalten kam nach diesen Zahlen nicht in Frage. Ganz löschen wollte ich es auch nicht. Für sehr große Ausgaben aus Playwright, langen Logs oder Web-Abrufen habe ich noch keine Messung. Dort richtet lean-ctx wenig aus.

Deshalb gibt es jetzt ein Startskript mit vier Zeilen. Es schaltet context-mode für genau eine Sitzung ein und lässt die globale Einstellung unberührt.

context-mode nur für eine sitzung

~/bin/claude-cm
#!/usr/bin/env bash
# context-mode nur für diese Sitzung (global aus).
# Für Aufgaben mit großem Tool-Output: Playwright, Logs, große Web-Fetches.
exec claude --settings '{"enabledPlugins":{"context-mode@context-mode":true}}' "$@"

Ob die Weiche greift, prüfst du am init-Event. Das ist die erste Meldung einer Sitzung im Format stream-json, sie listet alle geladenen Tools. Mit dem Skript waren es 11 Tools von context-mode, im normalen Start 0.

mit claude-cm11context-mode-tools im init-event
ohne claude-cm0context-mode-tools im init-event

prüfen, ob die weiche greift

terminal
claude-cm -p "ok" --model haiku --max-turns 1 \
  --output-format stream-json --verbose </dev/null \
  | grep -m1 '"subtype":"init"' \
  | grep -o 'mcp__plugin_context-mode[a-z_-]*' | sort -u | wc -l

vibesec: sicherheit beim schreiben statt danach

VibeSec ist eine einzige Markdown-Datei mit 758 Zeilen. Sie beschreibt Zugriffskontrolle, Eingabeprüfung, Uploads, SQL, Ausgabe-Kodierung, Weiterleitungen und den Umgang mit Secrets, jeweils mit Checklisten und typischen Fehlern. Ein Skript führt sie nicht aus. Das habe ich vor der Installation gelesen, denn ein Skill mit Shell-Zugriff wäre eine andere Entscheidung gewesen.

In meinem Ablauf gab es bisher zwei Sicherheitsprüfungen. Ein Hook liest jede Änderung, nachdem sie geschrieben ist. Bei Auth, Datenbank oder Infrastruktur folgt ein adversariales Review durch Codex. Beide finden Fehler, die schon im Code stehen. Jeder Fund kostet dann eine weitere Runde, und meine Grenze liegt bei drei.

VibeSec setzt vorher an. Bei Web-Code, der Anmeldung, Eingaben, Uploads, SQL, HTML-Ausgabe, Weiterleitungen, Secrets oder fremde Adressen berührt, lädt Claude Code den Skill in der BUILD-Phase. Der Code entsteht dann gleich mit Besitzprüfung auf der Datenebene und kodierter Ausgabe. Alles andere zahlt nichts, weil der Skill außerhalb dieser Fälle nicht geladen wird.

  1. define

    Ziel und Prüfbefehl.

  2. plan

    Schritte, Dateien, Risiken.

  3. build

    Web-Code mit Auth, Eingaben oder SQL: VibeSec zuerst laden.

    aktuell
  4. verify

    Prüfbefehle grün, Gegenprobe rot.

  5. review

    Security-Hook, danach Codex adversarial.

  6. ship

    Ship Gate, nur mit Freigabe.

Hook und Codex bleiben, wo sie sind. Das Review wird nicht überflüssig, es findet nur weniger. Jeder Befund, den Codex nicht mehr melden muss, spart eine Runde mit Diff, Befund und Nachbesserung.

stop-slop: installiert, aber ohne regel

stop-slop besteht aus sieben kleinen Dateien. Die Regeln sind gut. Sie verbieten Füllphrasen, Kontraste nach dem Muster „nicht X, sondern Y“ und Gedankenstriche und verlangen aktive Sätze mit konkreten Akteuren. Am Ende steht eine Bewertung über fünf Dimensionen mit einer Schwelle von 35 von 50 Punkten.

Fast genau diese Rubrik nutze ich schon. Für deutsche Texte läuft ein Humanizer als Pflicht, für englische ein zweiter Linter. Ein dritter Prüfer mit leicht anderen Wortlisten hätte zu jedem Text eine weitere Runde Abgleich gebracht. Deshalb liegt stop-slop im Skill-Ordner und kann per Aufruf genutzt werden, hängt aber an keiner Regel.

drei werkzeuge, drei entscheidungen

Drei Werkzeuge, drei verschiedene Entscheidungen
werkzeugversprichtgeprüftentscheidung
context-modebis zu 98 % weniger kontexta/b mit 6 läufen: +11 % kosten, gleiche antwortglobal aus, pro sitzung per claude-cm
VibeSecsicherer web-code aus sicht eines bug hunters758 zeilen regeln, nur markdown, kein skriptin BUILD bei web-code mit auth, eingaben, sql
stop-slopki-muster aus texten entfernen7 markdown-dateien, deckt sich mit zwei skillsinstalliert, in keine regel eingebaut

Die Tabelle wirkt uneinheitlich. Genau das ist das Ergebnis. Eine Pauschalregel wie „alles installieren“ oder „nichts Neues“ hätte bei mindestens einem der drei danebengelegen.

zwei regeln für die CLAUDE.md

~/.claude/CLAUDE.md
## Loop
- Token: lean-ctx bleibt Standard. context-mode global AUS (eigener A/B: +11 %).
  Aufgabe mit großem Tool-Output (Playwright, Logs, Web-Fetches > 50 KB)
  → eigene Sitzung per ~/bin/claude-cm.

## Security
- Web-Code mit Auth, Eingaben, Uploads, SQL, HTML-Ausgabe, Redirects,
  Secrets oder fremden URLs: in BUILD vorher vibesec-skill laden.
  Security-Audit: erst vibesec-skill, dann /codex:adversarial-review.

der september in vier messpunkten

Die Entscheidung von heute steht auf drei früheren Messungen. Jede davon hat ein Werkzeug aus dem Startkontext genommen oder draußen gehalten.

Vier Messpunkte im September: Sockel-Diät am 10.09., Headroom am 17.09., context-mode am 18.09., drei Skills am 29.09.

  1. 10.09. · erledigt

    sockel-diät

    Plugins und Skills aus dem Startkontext: 73.000 auf 17.000 Token.

  2. 17.09. · erledigt

    headroom-proxy

    0,7 % Kompression, Lauf 2,5 % teurer. Wieder ausgebaut.

  3. 18.09. · erledigt

    context-mode a/b

    Sechs Läufe, 11 % teurer. Abgeschaltet, installiert gelassen.

  4. 29.09. · in Arbeit

    drei skills

    VibeSec eingebaut, context-mode als Weiche, stop-slop geparkt.

so prüfst du ein neues werkzeug

1. erst lesen, dann installieren

Öffne jede Datei des Repos. Markdown ohne Skripte ist ein anderes Risiko als ein Plugin mit Hooks auf allen Tools.

2. einen a/b-lauf aufsetzen

Nimm eine echte Aufgabe aus deinem Alltag, keine Demo. Stelle beide Arme per --settings fest ein, damit kein globaler Schalter dazwischenfunkt. Vergleiche Kosten, cache_write und die Antwort.

3. die überschneidung suchen

Frag bei jedem Werkzeug, welches vorhandene schon dieselbe Arbeit macht. Doppelte Kompression und doppelte Textprüfung kosten, ohne etwas zu bringen.

4. den ort im ablauf festlegen

Ein Werkzeug ohne festen Platz wird entweder nie geladen oder immer. Schreib in deine CLAUDE.md, in welcher Phase und bei welchem Auslöser es greift.

Was jede zusätzliche Runde in Token und Geld bedeutet, rechnet „Was eine Aufgabe auf Opus 5.5 kostet“ vor. Wie ein Review an den geprüften Code-Stand gebunden wird, zeigt „ClaudeX Loop geprüft“.

was noch offen ist

Die Messung für context-mode bei wirklich großen Ausgaben steht aus. Sobald ein Playwright-Lauf mit mehreren Megabyte Ausgabe ansteht, läuft er zweimal, einmal mit Weiche und einmal ohne. Fällt der Unterschied klar zugunsten des Plugins aus, bekommt es einen festen Auslöser im Ablauf. Sonst bleibt es bei der Weiche.

häufige fragen

Spart context-mode in Claude Code Token?

In meinem Setup nicht. Im A/B-Test mit sechs Läufen derselben Aufgabe waren die beiden Läufe mit context-mode im Mittel rund 11 Prozent teurer, die Antwort war in allen Läufen richtig. lean-ctx komprimiert die Ausgaben bei mir schon vorher, context-mode legt dann vor allem zusätzliche Tools und Regeln in den festen Vorspann jeder Sitzung.

Wann lohnt sich context-mode trotzdem?

Bei Aufgaben mit sehr großer Tool-Ausgabe, etwa Playwright-Läufen, langen Logs oder großen Web-Abrufen. Das habe ich noch nicht gemessen. Deshalb starte ich solche Sitzungen gezielt mit einem kleinen Skript, das context-mode nur für diese eine Sitzung einschaltet.

Was macht der VibeSec-Skill?

VibeSec ist eine Markdown-Anleitung mit 758 Zeilen für sicheren Web-Code. Sie deckt Zugriffskontrolle, Eingabeprüfung, Uploads, SQL, Ausgabe-Kodierung, Weiterleitungen und Secrets ab. Der Skill führt keine Skripte aus, er ändert nur, wie Claude Code schreibt.

Ersetzt VibeSec ein Security-Review?

Nein. VibeSec wirkt beim Schreiben. Danach prüft bei mir ein Security-Hook jede Änderung, und bei sicherheitsrelevantem Code folgt ein adversariales Review durch Codex.

Warum ist stop-slop installiert, aber nicht eingebunden?

Die Regeln von stop-slop decken sich fast ganz mit zwei Skills, die ich schon für Texte nutze. Ein dritter Prüfer mit leicht anderen Regeln würde Texte nicht besser machen, nur die Prüfung länger.

Wie prüfe ich, ob ein Plugin in einer Sitzung aktiv ist?

Starte Claude Code mit -p, --output-format stream-json und --verbose. Das init-Event listet alle geladenen Tools. Mit dem Plugin tauchen dessen Tools auf, ohne das Plugin nicht. Bei mir waren es 11 gegen 0.

quellen und stand

Stand: 29.09.2026. Messwerte aus eigenen Läufen am 10., 17., 18. und 29.09.2026. Gelesen wurden die Repos von context-mode (Version 1.0.169), VibeSec und stop-slop.

A brass balance scale on a dark table in warm light. Two small weights sit in one pan, a larger one stands next to the scale on the table.
Title image AI-generated and labelled as such.

Three links landed on my desk today. context-mode promises 98 percent less context. VibeSec is meant to make Claude Code write web code like a bug hunter. stop-slop strips AI patterns from text. The request that came with them was clear: install all three and build them into the existing system, for better results and lower token use.

In the end each of the three tools got a different decision. VibeSec now sits firmly in the workflow. context-mode stays off and can be switched on per session. stop-slop is installed in the skills folder and is not tied to any rule.

This post belongs to the series about my Claude Code setup. "Setting up Claude Code" builds the base system, "The agent loop" describes the loop with three rounds, "Claude Code builds. Codex reviews" adds the review. You can read this one without the others.

the gist in four sentences

A tool that promises to save tokens has to compete in your own setup against a run without it. context-mode cost me 11 percent more, because lean-ctx already does the job. VibeSec only costs context when web code with a security angle is being written, and that is exactly where it belongs. For stop-slop there was no gap to fill.

context-mode in the a/b test+11 %more expensive per run, same answer
headroom-proxy0.7 %compression after lean-ctx
baseline per session17,000tokens, previously 73,000

the system before the update

My setup has a fixed flow from DEFINE to SHIP. The main model plans and approves the work. Cheaper instances build the individual steps. Codex reviews as an independent reviewer. Every session starts with lean-ctx, a local tool that compresses file reads, searches and shell output before they reach the context.

Until 10 September the fixed prefix of a session was around 73,000 tokens. After a clean-up that day it was 17,000. Since then one simple rule applies. Anything that wants into the start context has to save measurably more there than it costs.

The diagram shows where the three new tools attach. Tap a node to see its connections.

The system in four layers: session with lean-ctx and optional context-mode, orchestration with orchestrate and executor, review with VibeSec, security hook and Codex, writing with humanizer and stop-slop.

  • session: lean-ctx (Compresses reads, searches and shell output in every session. Default.); context-mode (Sandbox and search index for large output. Only via claude-cm, only for one session.)
  • orchestration: orchestrate (Plan with a check command per step, hard limit of three rounds.); executor (Cheaper instances build the steps, the main model accepts them.)
  • review: VibeSec (New: rules for secure web code, loaded while writing.); security-hook (Checks every change after it is written.); Codex (Independent review, adversarial for security topics.)
  • writing: humanizer (Mandatory for public German texts.); stop-slop (Newly installed, deliberately not tied to any rule.)

Verbindungen: von lean-ctx zu orchestrate, von context-mode zu executor, von orchestrate zu executor, von executor zu VibeSec, von VibeSec zu security-hook, von security-hook zu Codex, von executor zu humanizer.

Tap a node to see its connections.

context-mode: measured and found expensive

context-mode is a plugin with its own MCP server. It runs code in a sandbox, puts large output into a local search index and returns only the matching hits to the model. It also hooks into Read, Grep, WebFetch, Agent and every MCP tool. The idea is sound. Raw output eats most of the context in long sessions.

I had installed it back on 18 September. The A/B test ran the same day. The task was deliberately output-heavy. Claude had to find the three most-changed files from git log --stat of a repo with around 168 KB of output. Four runs without context-mode, two with it, each arm pinned via --settings.

cost per run in us dollars, the same task with context-mode off and on
CategoryValue
off 1$1.28
off 2$1.26
off 3$1.16
off 4$1.02
on 1$1.30
on 2$1.31
  • off 1$1.28
  • off 2$1.26
  • off 3$1.16
  • off 4$1.02
  • on 1$1.30
  • on 2$1.31

Without the plugin the runs cost between 1.02 and 1.28 dollars. With the plugin both runs came in at 1.30 and 1.31 dollars, on average around 11 percent more. All six answers were correct.

The cache shows why. With context-mode each session wrote around 116,000 tokens to the cache, without it 94,000 to 113,000. The plugin adds eleven tools and a block of routing rules to the prefix. The compression it offers had already been done by lean-ctx.

cache_write in thousand tokens: with context-mode the fixed prefix of every session grows
CategoryValue
off, min94k
off, max113k
on116k
  • off, min94k
  • off, max113k
  • on116k

There was also a trap. The context-mode hook replaces curl calls without -s -o file with an echo. A live check such as curl -sI on a cache header then silently returns nothing, and you go looking for the fault on the server.

A day earlier I had seen the same pattern with a compression proxy. After lean-ctx it compressed another 0.7 percent and made the run 2.5 percent more expensive. Two tools for the same job do not save twice. The second one only pays admission.

the fix: a switch per session

After these numbers, switching it on globally was out of the question. I did not want to delete it either. For very large output from Playwright, long logs or web fetches I have no measurement yet. That is where lean-ctx achieves little.

So there is now a four-line launcher. It turns context-mode on for exactly one session and leaves the global setting untouched.

context-mode for one session

~/bin/claude-cm
#!/usr/bin/env bash
# context-mode for this session only (off globally).
# For tasks with large tool output: Playwright, logs, large web fetches.
exec claude --settings '{"enabledPlugins":{"context-mode@context-mode":true}}' "$@"

You can check that the switch works in the init event. That is the first message of a session in stream-json format, and it lists every loaded tool. With the launcher there were 11 context-mode tools, with a normal start 0.

with claude-cm11context-mode tools in the init event
without claude-cm0context-mode tools in the init event

check that the switch works

terminal
claude-cm -p "ok" --model haiku --max-turns 1 \
  --output-format stream-json --verbose </dev/null \
  | grep -m1 '"subtype":"init"' \
  | grep -o 'mcp__plugin_context-mode[a-z_-]*' | sort -u | wc -l

vibesec: security while writing, not afterwards

VibeSec is a single Markdown file with 758 lines. It covers access control, input validation, uploads, SQL, output encoding, redirects and handling secrets, each with checklists and typical mistakes. It runs no script. I read it before installing, because a skill with shell access would have been a different decision.

My workflow had two security checks so far. A hook reads every change after it is written. For auth, databases or infrastructure an adversarial review by Codex follows. Both find flaws that are already in the code. Each finding then costs another round, and my limit is three.

VibeSec comes in earlier. For web code that touches login, input, uploads, SQL, HTML output, redirects, secrets or external URLs, Claude Code loads the skill in the BUILD phase. The code is then written with ownership checks at the data layer and encoded output from the start. Everything else pays nothing, because the skill is not loaded outside these cases.

  1. define

    Goal and check command.

  2. plan

    Steps, files, risks.

  3. build

    Web code with auth, input or SQL: load VibeSec first.

    current
  4. verify

    Checks green, counter-check red.

  5. review

    Security hook, then an adversarial Codex review.

  6. ship

    Ship gate, only with approval.

Hook and Codex stay where they are. The review does not become redundant. It just finds fewer issues. Every finding Codex no longer has to report saves a round of diff, finding and fix.

stop-slop: installed, but without a rule

stop-slop consists of seven small files. The rules are good. They ban filler phrases, contrasts of the "not X, but Y" kind and em dashes, and they ask for active sentences with concrete actors. At the end there is a rating across five dimensions with a threshold of 35 out of 50 points.

I already use almost exactly this rubric. A humanizer is mandatory for German texts, and a second linter checks English ones. A third checker with slightly different word lists would have added another round of reconciliation to every text. So stop-slop sits in the skills folder and can be called on demand, but it is not tied to any rule.

three tools, three decisions

Three tools, three different decisions
toolpromisestesteddecision
context-modeup to 98 % less contexta/b with 6 runs: +11 % cost, same answeroff globally, per session via claude-cm
VibeSecsecure web code from a bug hunter's view758 lines of rules, markdown only, no scriptin BUILD for web code with auth, input, sql
stop-slopremove ai patterns from text7 markdown files, overlaps with two skillsinstalled, not tied to any rule

The table looks inconsistent. That is exactly the result. A blanket rule like "install everything" or "nothing new" would have been wrong for at least one of the three.

two rules for CLAUDE.md

~/.claude/CLAUDE.md
## Loop
- Tokens: lean-ctx stays the default. context-mode globally OFF (own A/B: +11 %).
  Task with large tool output (Playwright, logs, web fetches > 50 KB)
  → separate session via ~/bin/claude-cm.

## Security
- Web code touching auth, input, uploads, SQL, HTML output, redirects,
  secrets or external URLs: load vibesec-skill in BUILD first.
  Security audit: vibesec-skill first, then /codex:adversarial-review.

september in four measurements

Today's decision rests on three earlier measurements. Each of them took a tool out of the start context or kept it out.

Four measurements in September: baseline diet on Sep 10, Headroom on Sep 17, context-mode on Sep 18, three skills on Sep 29.

  1. sep 10 · done

    baseline diet

    Plugins and skills out of the start context: 73,000 down to 17,000 tokens.

  2. sep 17 · done

    headroom-proxy

    0.7 % compression, run 2.5 % more expensive. Removed again.

  3. sep 18 · done

    context-mode a/b

    Six runs, 11 % more expensive. Switched off, kept installed.

  4. sep 29 · in progress

    three skills

    VibeSec built in, context-mode as a switch, stop-slop parked.

how to test a new tool

1. read first, then install

Open every file in the repo. Markdown without scripts is a different risk from a plugin with hooks on every tool.

2. set up an a/b run

Take a real task from your daily work, not a demo. Pin both arms via --settings, so no global switch gets in the way. Compare cost, cache_write and the answer.

3. look for the overlap

Ask of every tool which existing one already does the same job. Double compression and double text checks cost without adding anything.

4. fix its place in the workflow

A tool without a fixed place is either never loaded or always. Write into your CLAUDE.md in which phase and on which trigger it applies.

What each extra round means in tokens and money is worked out in "What a task costs on Opus 5.5". How a review is bound to the reviewed code state is shown in "ClaudeX Loop reviewed".

what is still open

The measurement for context-mode with really large output is still pending. As soon as a Playwright run with several megabytes of output comes up, it runs twice, once with the switch and once without. If the difference clearly favours the plugin, it gets a fixed trigger in the workflow. Otherwise the switch stays.

frequently asked questions

Does context-mode save tokens in Claude Code?

Not in my setup. In an A/B test with six runs of the same task, the two runs with context-mode were on average around 11 percent more expensive, and the answer was correct in every run. lean-ctx already compresses the output beforehand, so context-mode mainly adds extra tools and rules to the fixed prefix of each session.

When is context-mode still worth it?

For tasks with very large tool output, such as Playwright runs, long logs or large web fetches. I have not measured that yet. That is why I start such sessions deliberately with a small script that turns context-mode on for that one session only.

What does the VibeSec skill do?

VibeSec is a Markdown guide with 758 lines for secure web code. It covers access control, input validation, uploads, SQL, output encoding, redirects and secrets. The skill runs no scripts. It only changes how Claude Code writes.

Does VibeSec replace a security review?

No. VibeSec works while writing. After that a security hook checks every change in my setup, and security-relevant code gets an adversarial review by Codex.

Why is stop-slop installed but not wired in?

The rules of stop-slop overlap almost completely with two skills I already use for text. A third checker with slightly different rules would not make texts better, only the check longer.

How do I check whether a plugin is active in a session?

Start Claude Code with -p, --output-format stream-json and --verbose. The init event lists every loaded tool. With the plugin its tools show up, without it they do not. In my case it was 11 against 0.

sources and notes

As of 29 September 2026. Measurements from my own runs on 10, 17, 18 and 29 September 2026. I read the repos of context-mode (version 1.0.169), VibeSec and stop-slop.

geschrieben vonwritten by · business transformation, ki und führung im mittelstand. methode: core+.business transformation, ai and leadership for mid-sized companies. method: core+.

weiterlesenkeep reading

← zurück zum herrlichblog← back to the herrlichblog