Lesson 15 / 26
Testing Triggering With Should and Should-Not Prompts
Build a small list of requests that must and must not use the skill.
Both false negatives and false alarms cost you
A description can fail in two ways: a false negative (the skill does not trigger when it should, so the user gets no help) and a false positive (it triggers on unrelated requests, wasting context and sometimes causing wrong behaviour). Test both. Write 10 to 20 realistic prompts: some that should trigger (varied wording, casual and formal), some that should not (similar-sounding but different jobs). Run them in a fresh session and record whether the skill was used; adjust the description's keywords and "when" clause until the results are right, then keep the list as a regression suite and re-run it whenever you change the description or add a similar skill. The exact decision is made by Claude, so you test with the real tool, but a cheap word-overlap check like the one below can catch descriptions that share no vocabulary with the prompts you expect.
Test the triggers, the scripts, the result
Check that the skill triggers when it should, that its scripts work, and that outputs are right; then improve from real transcripts.
A trigger test list with a crude check, run
I ran this with plain Python 3 (standard library only). This is a simple simulation of the idea, not how Claude actually decides: Claude uses its own judgement over the descriptions, which is why clear descriptions matter. Three prompts that should trigger the release-notes skill all share keywords with it, and three unrelated prompts share none, so there are no false alarms in this crude check. A real test runs the prompts in the actual tool and records whether the skill was used.
SHOULD_TRIGGER = ["write the changelog for this release", "draft release notes", "what changed since v1.2"]
SHOULD_NOT = ["fix the failing test", "explain this regex", "write a haiku"]
TRIGGER_WORDS = {"changelog", "release", "notes", "changed", "tag"}
def would_trigger(prompt): return len(set(prompt.lower().replace("?", "").split()) & TRIGGER_WORDS) > 0
tp = sum(would_trigger(p) for p in SHOULD_TRIGGER); fp = sum(would_trigger(p) for p in SHOULD_NOT)
print(f"should trigger : {tp}/{len(SHOULD_TRIGGER)} triggered")
print(f"should NOT : {fp}/{len(SHOULD_NOT)} triggered (false alarms)")
for p in SHOULD_TRIGGER + SHOULD_NOT: print(f" {'TRIGGER' if would_trigger(p) else ' - '} {p}")
Output:
should trigger : 3/3 triggered
should NOT : 0/3 triggered (false alarms)
TRIGGER write the changelog for this release
TRIGGER draft release notes
TRIGGER what changed since v1.2
- fix the failing test
- explain this regex
- write a haikuQuick check: Why include prompts that should NOT trigger the skill?
- To detect false positives that waste context or cause wrong behaviour
- To make the list longer
- Because skills must fail sometimes
- There is no reason
Answer
To detect false positives that waste context or cause wrong behaviour — A skill that triggers on everything is as broken as one that triggers on nothing.