Publish, measure, improve
Build plugins people keep using.
Track plugin usage, reported outcomes, and evaluation results by version. Run evals with Promptfoo, register releases with Telvine, and keep the evidence behind your release decisions.
npm i -g @telvine/cli telvine login telvine init telvine publish ./my-plugin
Why plugin quality matters
Successful use gives people a reason to return
“We look at retention numbers. We look at, you know, how successful the plugin, the quality.”
Speaking about OpenAI’s plugin recommendations, he explained that poor plugins can stop being recommended. For builders across ChatGPT, Claude, and other runtimes, the practical question is the same: does the plugin complete useful work, and do people come back?
Read the guide to plugin quality and discovery →Versioning and promotion
Review the evidence before your next release
Know whether your next release is ready. Telvine brings evaluation results, production outcomes, and user feedback together for each plugin version.
Run evaluations in your own harness with Promptfoo. Telvine checks required coverage, grader results, and the suite pass threshold before stable promotion. Review production signals alongside those checks.
Illustrative evidence · production signals require review
Example dashboard
See how your plugin performs
Track usage, errors, latency, and feedback in one place. Compare releases to see what improved and where to investigate. Illustrative data shown below.
What is an Agent Plugin?
Plugins are the apps users install in their agents. A plugin bundles capabilities such as Skills, tools, connectors, and hooks. Skills are task-specific instructions inside a plugin.
Telvine registers your plugin inventory, measures observable components, and helps you review each release. Use your existing agents and models, and forward events to your analytics stack.
Publish, measure, and improve your plugins
Measure plugin performance
Track reported outcomes, latency, component usage, and feedback wherever a runtime integration can emit Telvine events.
Evaluate quality
Run native Promptfoo suites locally or in CI. Record assertion results and repeated trials against a fixed suite revision and plugin version.
Compare and release versions
Review eval coverage, weighted pass rates, and regressions by version. Telvine blocks stable promotion when required evaluation evidence fails the release policy.
Keep your stack
Forward normalized events to PostHog, Mixpanel, Amplitude, Datadog, your warehouse, or internal dashboards while Telvine manages the plugin improvement record.
Who Telvine is for
SaaS teams shipping MCP servers and agent plugins who need to prove their plugins work across many agent surfaces. Platform and AI teams standardizing packaging, rollout, telemetry, evals, and promotion gates across an org. Agencies and consultancies proving whether the plugins they build actually improve in production.
Pricing
Start free. Scale with your plugins.
Choose a plan based on your tracked plugins, event volume, and team size. Every plan includes the CLI. Promptfoo execution runs in your environment; model-provider charges are separate. Contact us for paid-plan pricing and confirmed limits.
Free
For trying Telvine on one plugin.
- 1 tracked plugin
- Up to 3 tracked components
- Metadata events for local and design-partner testing
- 1 dashboard user
- CLI included
Builder
For builders shipping their first plugin portfolio.
- 3-10 tracked plugins
- Plugin portfolio dashboards
- Higher metadata event allowance
- Eval suites and version comparison workflows
- 1-3 dashboard users
- CLI included
Team
For teams operating plugins across customers or departments.
- 10+ tracked plugins
- Production event allowance with overage path
- Shared eval, feedback, and promotion workflows
- Webhook and CSV export workflows
- 5-10 dashboard users
- Extra seats available
Enterprise
For organizations with security, scale, and procurement needs.
- Contracted plugin and event volume
- Custom retention and data handling review
- SSO, audit, and governance requirements reviewed during scoping
- Security and procurement review
- Dedicated support
Learn the playbook
Guides across agent platforms on building, measuring, evaluating, and improving Agent Plugins and their Skills.
- What is an Agent Plugin?
- How to build an Agent Plugin
- How to create a ChatGPT plugin
- How to create a ChatGPT desktop plugin
- How to add a plugin marketplace to ChatGPT
- How to build a plugin for OpenAI Codex
- How to build a plugin for Claude Code and Claude Cowork
- How to add a plugin marketplace to Codex
- Forkable Agent Plugin templates
- What is an agent Skill?
- How to create an agent Skill
- How agent plugins get discovered and recommended
- Run Promptfoo evals and record release evidence
- How to measure Agent Plugin usage
- Browse all guides →
Frequently asked questions
- What is an Agent Plugin?
- An Agent Plugin is an installable package of capabilities for an agent harness. It can contain Skills, tools, connectors, hooks, or agent configuration. Plugins are the apps people install; Skills are task-specific capabilities inside them.
- What is Telvine?
- Telvine helps teams publish agent plugins, measure performance, evaluate quality, and compare versions before releasing an update. Run evaluations with Promptfoo or your existing evaluator; Telvine keeps the plugin version and release evidence.
- Which agents and runtimes does Telvine support?
- Telvine tracks plugin versions and evaluation evidence across ChatGPT, Codex, Microsoft 365 Copilot Cowork, Claude Cowork, Claude Code, and other agent runtimes. Runtime telemetry requires an integration or wrapper that can emit Telvine events. Available components and execution behavior vary by surface; test each environment you intend to support.
- How do I instrument a Plugin?
- Publish the plugin directory with the CLI, import its evals, and emit metadata-only component events from observable runtime boundaries. Telvine captures invocation metadata, errors, feedback, outcomes, latency, and version signals while your existing analytics stack remains the system of record.
- Can my team define plugin evaluations?
- Yes. PM-authored eval suites can live with the plugin as versioned assets today, and Telvine treats evals as part of the improvement loop: define synthetic scenarios, expected outcomes, versioned graders, and release policies. Promptfoo executes and grades locally; Telvine records metadata-only results.
- How does Telvine help plugin owners ship better output?
- Each plugin version carries its component inventory, eval results, production outcomes, feedback, latency, and error evidence. Owners can review results by version. Required eval coverage, assertions, and pass thresholds must clear the configured policy before Telvine marks a release stable. Distribution remains in your marketplace or repository.
- Does Telvine collect prompts or user content?
- Telvine is designed for metadata-only telemetry. Keep prompts, outputs, customer records, and tool payloads out of event fields and optional feedback. Promptfoo execution evidence stays in your controlled environment; the Telvine adapter submits only allowlisted result metadata.
- Why does this matter now?
- Teams are shipping plugins into agent harnesses without a reliable way to know whether a new version actually helped. Existing tools can store events; Telvine gives teams the plugin performance data, eval signals, and version evidence to improve what agents do.
- Does Telvine replace PostHog, Mixpanel, Amplitude, or my warehouse?
- Telvine keeps the plugin version and evaluation record. Your analytics stack continues to handle broader product analytics; forward or export Telvine metadata events into the tools your team already uses.
Ship your next plugin version with confidence.
Start with one plugin. Measure what happens, compare versions, and improve your next release.