Telvine Learn

Claude Code Plugin Evaluation: Test Triggers Before Release

Evaluate Claude Code plugins with positive, negative, and paraphrased trigger tests. Use Telvine to version cases and record evaluation results.

Claude Code plugin evaluation starts with a concrete question: does the right component activate for the right task? Test both activation and the completed workflow before shipping a new plugin version.

Use your Claude Code test harness to execute the plugin. Use Telvine to inventory the package, import versioned eval cases, and record results alongside production telemetry. Telvine's current CLI imports suites; it does not launch Claude Code or generate and execute trigger tests for you.

What to test in a Claude Code plugin

A plugin is the installable product. Skills, agents, commands, hooks, and MCP configuration are components with different activation mechanisms. Write expectations for each mechanism rather than assuming every component should activate from natural language.

ComponentTest questionEvidence to inspect locally
SkillDoes the capability activate for an in-scope request?Harness activation trace and workflow result
AgentIs the intended agent selected or delegated to?Harness delegation trace
CommandDoes explicit invocation select the expected command?Command execution record
HookDoes the configured lifecycle event run the hook?Hook execution record
MCP toolDoes the workflow call the expected available tool?Local tool-call trace

Component activation and task success are separate checks. A Skill can activate correctly and still produce the wrong result.

Build a small trigger test matrix

For a synthetic invoice-review plugin, start with these cases:

Case typeTest requestExpected behavior
PositiveReview these synthetic invoices for overdue payments.Invoice-review Skill activates and identifies overdue items.
ParaphraseWhich test invoices need chasing?The same Skill activates despite different wording.
NegativeExplain what an invoice is.Invoice-review Skill stays inactive.
BoundaryReview invoices and send reminders.Review is permitted; sending follows the plugin's declared authorization rules.
OverlapSummarize this synthetic cashflow report.The appropriate reporting capability is selected.

Run each case in a fresh session with the same plugin version, model, harness configuration, and fixture set. Repeat cases to reveal inconsistent activation. Record those settings so a baseline and candidate comparison is meaningful.

Register the evaluation with Telvine

Put synthetic test scenarios under evals/invoice-triggers/cases.jsonl in your plugin repository. The Telvine evaluation walkthrough includes a copyable case file and the result-recording API workflow.

npm i -g @telvine/cli
telvine login
telvine publish ./my-plugin --dry-run
telvine publish ./my-plugin

The dry run lists discovered components and eval suites. Publishing registers the package and imports the cases. Install the plugin separately in your Claude Code test environment and execute the scenarios there.

Judge the results

Prefer an observed component activation signal over an agent's statement that it used a Skill. Mark a case as unobservable when your harness cannot expose the required evidence; do not turn a guess into a pass.

Report positive activation rate, negative false activations, and task success separately. A single overall pass rate can hide a plugin that activates on every request. Keep an explicit release rule in the suite's README.md, such as requiring all negative cases to pass and reviewing every changed activation before release. Telvine imports that text as the promotion-gate description; the text itself does not enforce a runtime gate.

Inspect full traces in your controlled test environment. Submit only scores, statuses, durations, and other safe result metadata to Telvine. Do not send user prompts, file contents, connector payloads, tool arguments, or model outputs.

Frequently asked questions

Does Telvine run Claude Code plugin tests automatically?

The current CLI publishes plugin inventory and imports eval cases. Your harness executes the tests, and an integration submits results through the Telvine API. Publishing a suite does not execute it.

Is static validation enough to evaluate a plugin?

Static checks catch package and instruction problems. Runtime trigger tests check actual activation, and workflow tests check whether the task succeeds. Use all three for release confidence.

Can I use an existing plugin evaluator?

Yes. Keep your evaluator for test execution and adapt its case results to Telvine's result schema. Preserve the model, harness, plugin version, and case identity when recording each run.

Continue with Skill trigger testing, the Telvine evaluation walkthrough, or Claude plugin packaging.