How to Test Skill Triggering: Positive and Negative Evals
Test whether a SKILL.md capability activates correctly, catches paraphrases, and avoids false triggers. Track versioned test cases and results with Telvine.
To test Skill triggering, present realistic requests to your target harness and observe which capability it activates. Include requests that should activate the Skill and requests that should leave it inactive. A successful example alone does not show that the trigger boundary works.
Define the Skill boundary first
Write down the task the Skill handles, the inputs it needs, and closely related tasks it should leave to other capabilities. For an invoice-review Skill, reviewing overdue synthetic invoices is in scope; explaining accounting terminology is outside its execution scope.
The SKILL.md description should make that distinction understandable to the harness. See how to create an agent Skill for the capability structure.
Create four groups of requests
- Positive: direct requests for the intended task.
- Paraphrased: the same task expressed with different vocabulary or sentence structure.
- Negative: related words without a request for the capability's task.
- Ambiguous or overlapping: requests that could select another installed Skill or need clarification.
For example, “Which synthetic invoices are overdue?” and “Who needs chasing in this test invoice list?” should select an invoice-review capability. “What does overdue mean?” should not. State the expected behavior before running the test so the observed outcome cannot redefine success.
Observe activation in the target harness
Run each request in a clean session, with a fixed set of installed plugins and the same model settings. Use a harness trace that exposes capability activation when available. An output that resembles the Skill's answer is weaker evidence than an observed activation.
If your harness cannot reveal Skill selection, label that check unobservable and assess task success separately. Repeat requests and retain each trial's result: activation can vary between runs.
Calculate useful trigger metrics
| Metric | Calculation | What it reveals |
|---|---|---|
| Positive activation rate | Correct activations / requests that should activate | Missed triggers |
| Negative false activation rate | Incorrect activations / requests that should stay inactive | An overbroad description |
| Paraphrase activation rate | Correct activations / paraphrased positive requests | Dependence on exact wording |
| Task success rate | Successful workflows / attempted workflows | Whether activation leads to useful work |
For an illustrative test with 20 positive requests and 20 negative requests, 18 correct positive activations means a 90% positive activation rate. Two activations on negative requests means a 10% false activation rate. These are example calculations, not Telvine benchmarks.
Version the cases and compare changes with Telvine
Store synthetic scenarios in evals/skill-triggers/cases.jsonl inside the plugin. Publish that plugin with Telvine to import the suite and record its component inventory:
npm i -g @telvine/cli
telvine login
telvine publish ./my-plugin --dry-run
telvine publish ./my-plugin
Run the cases in your harness, then submit case results through the Telvine API. The step-by-step walkthrough explains how to create a run, attach results, and complete it. Telvine stores the evidence; your harness performs activation detection and test execution.
Change one part of the description at a time and rerun the same cases against the candidate. Review positive and negative groups separately before deciding whether the new description is better.
Diagnose common trigger failures
| Failure | First thing to inspect |
|---|---|
| Direct request misses the Skill | Whether the description names the user's task clearly |
| Only exact phrasing works | Missing synonyms and realistic task language |
| Related questions trigger unnecessarily | An unclear scope boundary |
| Another Skill is selected | Overlap between installed capability descriptions |
| Activation works but the task fails | Instructions, tools, fixtures, and output checks |
Frequently asked questions
Should every request mentioning a keyword trigger the Skill?
No. The requested task determines whether activation is appropriate. Negative cases should include the same vocabulary as positive cases to test that distinction.
Does an invocation event prove the Skill triggered correctly?
It proves an invocation was observed. Correctness requires matching that observation against the case's expected activation and checking the workflow outcome.
Can I upload real user requests as test cases?
Use synthetic scenarios in published eval cases. Production telemetry must exclude prompts, file contents, connector payloads, tool arguments, and model outputs. Keep sensitive traces in your own controlled environment.
See Claude Code plugin evaluation for component-specific tests and plugin usage measurement for production signals.