Telvine Learn

How to Measure and Improve Agent Plugin and Skill Usage

Learn how to measure and improve Agent Plugins and Skills: package adoption, Skill invocation, component events, evals, version comparison, and privacy-safe telemetry.

Once you've built an Agent Plugin, the next question is whether it actually works and whether the next version improved. This guide covers how to measure and improve Agent Plugin and Skill usage - package adoption, Skill execution, component behavior, evals, feedback, and version comparison - so you can improve a plugin with evidence instead of guesswork.

The metrics that matter

An Agent Plugin is software, so measure it at two layers: the installed package and the capabilities inside it.

  • Observed plugin instances - how many installation ids have emitted a plugin event.
  • Plugin activation and retention - do observed instances keep using any capability?
  • Component inventory - which Skills, manifests, commands, MCP servers, hooks, connectors, apps, agents, directories, and runtime components shipped in each version.
  • Skill invocations - how often each capability actually runs.
  • Error rate - share of runs or component operations that fail.
  • Latency - p50 / p95 duration; the p95 is where pain hides.
  • Outcome mix - completed vs. partial vs. blocked, not just "did it end."
  • Autonomy rate - how often it ran with no human step-in.
  • Downstream actions - the real effects it triggered, such as a report exported or a ticket created.
  • Eval pass rate - how PM-authored scenarios and rubrics score each plugin or Skill version.

Measure successful use and retention together

An invocation records activity. A successful outcome records whether the plugin did the promised job. Retention records whether an observed instance came back. Examine all three across ChatGPT, Claude, and other runtimes using the same definitions, then separate results by runtime and version.

For example, a weekly invoice-exception review might produce these illustrative figures, calculated from observable metadata:

MetricStable cohortCandidate cohort
Instances first observed in week 1100100
Instances with a successful review in week 16075
Activated instances with another successful review in week 22436
Activation: successful first review / first observed instances60%75%
Successful repeat use: week 2 successful instances / week 1 activated instances40%48%

Count distinct installation ids, not total calls. Keep each cohort anchored to its first observed week and give both cohorts the full follow-up window. Report ordinary return activity separately from successful repeat use; an error retry is activity but may not represent delivered value.

These figures suggest a candidate worth investigating. They do not establish that the release caused the change: compare sample sizes, uncertainty, runtime mix, and workflow differences before promoting. Label instances that change version during follow-up so they do not silently contaminate the comparison.

For a monthly close plugin, use a monthly observation window. For a one-off migration, completion and quality may be more useful than weekly retention. Background runs should be distinguished from user-initiated use wherever the integration can observe that distinction.

Use existing skill.* and plugin.component.* events with permitted outcome metadata. Calculate cohorts in Telvine or your analytics stack where those signals are available. An observed installation id identifies an instance, not necessarily a unique person or a verified marketplace install.

See how plugin quality relates to discovery.

The events to emit

Capture the plugin and capability lifecycle as a small set of typed events:

  • plugin.install / plugin.update.applied - package adoption and version changes when the host, wrapper, or first-run telemetry can observe them.
  • /v1/plugins/:id/versions - declared component inventory for a shipped version.
  • skill.invocation.start / skill.invocation.end - a Skill run begins and finishes.
  • skill.invocation.error - a Skill run fails with an error class.
  • plugin.component.invoked / plugin.component.error - a connector, hook, MCP wrapper, app, agent, or package directory emits observable behavior.
  • feedback.submitted - a user rates the result.

Crucially, these should be a closed envelope of metadata - counts, durations, enums, ratings, component directories, and plugin-relative package paths - and never carry prompts, file contents, connector payloads, tool arguments, retrieved records, absolute local paths, or user file paths. You can measure everything above without touching sensitive data.

Treat marketplace installs as unverified until observed

Most agent harnesses do not yet expose reliable marketplace-install webhooks. Clicking install in ChatGPT, Codex, Claude Cowork, or another harness may not call your analytics endpoint. The durable pattern is first-run telemetry:

  1. On first plugin execution, create or read a stable installation id.
  2. Emit plugin.install once for that id.
  3. Emit skill.invocation.* and plugin.component.* events for each observed capability or component operation.

If plugin.install is unavailable, count the first committed runtime event as an observed instance, but do not present it as a verified marketplace install.

Measure directories without collecting content

Every plugin version should register its component inventory. For Claude, that means mapping directories and manifest fields such as skills/, commands/, agents/, hooks/, .mcp.json, .lsp.json, output-styles/, themes/, monitors/, bin/, scripts/, and settings.json.

Use plugin-relative metadata:

{
  "component_type": "directory",
  "name": "monitors",
  "telemetry_mode": "component_events",
  "component_path": "monitors/monitors.json",
  "component_directory": "monitors",
  "manifest_field": "experimental.monitors",
  "source": "default_directory",
  "experimental": true
}

Then emit plugin.component.invoked or plugin.component.error when your build, export, wrapper, hook, or runtime can observe a directory-level operation such as scanned, validated, resolved, cached, or installed. Inventory is not usage: if the host does not expose a signal, keep the component as declaration_only or host_unavailable.

Compare versions, don't just track totals

The most useful question is "did my new plugin version help?" Answer it by comparing two versions head-to-head on adoption, invocation rate, error rate, latency, and rating - with statistical significance, so a noisy 2% wiggle isn't mistaken for a win. A good comparison tells you one of three things: it improved, it regressed, or nothing changed.

Watch for tradeoffs, too: a release can cut errors but get slower. Looking at one metric at a time hides that.

Add evals to the improvement loop

Production traces tell you where users actually struggle. Evals tell you whether a proposed fix improves known scenarios before you promote it. Keep PM-authored eval suites with the plugin repo:

my-plugin/
  skills/expense-readiness-review/SKILL.md
  evals/expense-readiness-review/eval.yaml
  evals/expense-readiness-review/cases.jsonl

An eval suite should define the scenario, expected outcome, rubric criteria, pass threshold, blocker cases, fixture pointers, and promotion criteria. Keep it metadata-safe: no live customer data, production prompts, connector payloads, browser DOM, screenshots, account values, or customer content.

Find where failures concentrate

A single error rate is an average. To fix the right thing, break failures down by reason and look at the concentration — usually a small number of causes explain most failures (a Pareto pattern). Fix those first.

How to do this with Telvine

Telvine is built for exactly this. Publish your plugin with the CLI, register its component inventory, and add first-run telemetry where the host does not provide install webhooks. You get observed plugin instances, Skill invocations, errors, and feedback, plus:

  • Version comparison with significance testing.
  • PM-authored eval suites that complement production traces and feedback.
  • Funnels, retention, latency percentiles, outcome and autonomy breakdowns.
  • Per-component, per-directory, and per-script reliability.
  • Webhooks and CSV export to pipe events into your warehouse, PostHog, Mixpanel, Amplitude, or internal dashboards.

No dashboards to build, no user content collected.

Frequently asked questions

Can I measure plugins without sending data to a third party? Yes — emit the typed events to your own endpoint, or use webhooks and export to fan them into your existing stack.

What's the single most useful metric? Start with successful completion of the plugin’s promised job. Pair it with repeat successful use at the workflow’s natural cadence, then inspect errors and latency to understand where people struggle.

Does measuring require code changes to my Skill? No - publishing the plugin and wrapping the Skill adds instrumentation around it; the Skill logic stays the same.

Next steps