Telvine Learn

How to Evaluate Agent Plugins with Telvine

Import synthetic plugin eval cases, run them in your own harness, and record versioned results with Telvine's evaluation API.

Use Telvine to keep plugin eval cases, component inventory, and run results attached to the versions you are testing. Your harness executes the scenarios and judges the outcome; Telvine imports the suite and stores the evaluation evidence.

This walkthrough covers the current publish CLI and evaluation API. There is no telvine eval command in the current CLI.

1. Add a synthetic trigger suite

Start with a plugin containing a supported manifest and a Skill:

my-plugin/
  .claude-plugin/plugin.json
  skills/invoice-review/SKILL.md
  evals/invoice-triggers/cases.jsonl
  evals/invoice-triggers/rubric.md
  evals/invoice-triggers/README.md

For a Codex install surface, use .codex-plugin/plugin.json. See the plugin build guide for packaging.

Put each case on a separate line in cases.jsonl:

{"case_key":"positive-overdue","name":"Review overdue invoices","scenario":"Review the synthetic invoice fixture for overdue payments.","expected_outcome":"Invoice-review Skill activates and identifies the overdue invoices.","tags":["positive"],"weight":1}
{"case_key":"negative-definition","name":"Explain invoice terminology","scenario":"Explain what an overdue invoice means without reviewing records.","expected_outcome":"Invoice-review Skill does not activate.","tags":["negative"],"weight":1}
{"case_key":"paraphrase-chasing","name":"Find invoices to chase","scenario":"Which invoices in the synthetic test fixture need chasing?","expected_outcome":"Invoice-review Skill activates and identifies the overdue invoices.","tags":["paraphrase"],"weight":1}

These scenarios are synthetic authored test content. Do not paste production prompts or customer records into the suite. Keep fixture files and full execution evidence in your controlled test environment.

In rubric.md, describe the checks your evaluator should apply: observed activation, expected task outcome, and any permission boundary. The CLI imports a hash of this rubric text; your evaluator must read and apply the local rubric itself.

In README.md, state the release decision rule, such as “all negative cases must pass; review every regression.” The CLI imports this as the suite's promotion-gate description. The publish importer currently uses a default pass threshold of 0.8; that overall threshold does not replace your per-group review or automatically enforce the rule.

2. Inspect and publish

npm i -g @telvine/cli
telvine login
telvine publish ./my-plugin --dry-run
telvine publish ./my-plugin

Confirm the dry run lists the expected Skill and eval suite. Publishing records the component inventory, imports evals/**/cases.jsonl, and prints the imported suite ID. It does not execute the tests or install the plugin into your harness.

3. Execute the cases in your own harness

Install the plugin version into the target test environment. Run each case in a fresh session with fixed model settings and fixtures. Record activation observations and outcome checks locally.

Test both the stable version and the candidate with the same case set. Keep trial results separate when repeating a case: the API stores one result per case per run, so separate trials should use separate runs.

For trigger-test design, use the Skill triggering guide or the Claude Code component test matrix.

4. Record the run through the API

Use an authenticated Telvine API client with access to the plugin's organization. Read operations need read access; run and result writes need write access. Keep credentials in the client configuration rather than in case files. Authenticate requests with an Authorization: Bearer TOKEN header. A valid Clerk session can access both operations; API keys must match the requested scope, so use a read key for retrieval and a write key for run/result submission, each scoped to the plugin or its organization.

First request GET /v1/eval-suites/SUITE_ID to retrieve the imported cases and their IDs. Then send POST /v1/eval-suites/SUITE_ID/runs with a body like this, replacing the example versions and model with the values actually tested:

{
  "plugin_version": "0.2.0",
  "baseline_version": "0.1.0",
  "harness": "claude-code",
  "model": "your-tested-model",
  "status": "running"
}

Use the returned run ID to send POST /v1/eval-runs/RUN_ID/cases for each result:

{
  "eval_case_id": "IMPORTED_CASE_ID",
  "status": "passed",
  "score": 1,
  "rubric_results": {
    "activation_matched": true,
    "outcome_matched": true
  },
  "duration_ms": 4200
}

Only set passed after applying the case's expected behavior and rubric. For a negative case, staying inactive is the expected result. Supported case statuses are passed, failed, skipped, and errored. Use a skipped result for a check you cannot observe, and do not treat it as proof of correctness.

Keep submitted results metadata-only. Do not place transcripts, outputs, tool arguments, or customer content in rubric_results, feedback, or evidence references. A safe evidence reference can be an opaque identifier for evidence retained in your environment.

Complete the run with POST /v1/eval-runs/RUN_ID/complete:

{"status":"completed"}

Use GET /v1/eval-runs/RUN_ID to inspect the stored results, and GET /v1/plugins/PLUGIN_ID/evals/summary to read the plugin's suite summaries. Run completion means recording has finished; it does not mean every case passed.

5. Review the release evidence

Compare failed cases and trigger groups as well as the aggregate pass rate. Review skipped and errored checks, hold the candidate when required evidence is missing, and make the release decision using the suite's stated rule.

After release, measure plugin usage with skill.* events for Skills and plugin.component.invoked or plugin.component.error for non-Skill behavior. Production invocations and error rates complement the synthetic tests; they do not by themselves prove correct trigger selection.

Frequently asked questions

Can Telvine generate and execute the cases for me?

The current CLI imports cases you supply. Use your existing harness or evaluator for generation, execution, and judging, then map its results to the API schema above.

Does publishing send fixture files or runtime transcripts?

The eval importer sends the authored scenario and expected outcome, along with case metadata. It does not upload fixture files or runtime transcripts through this import path. Keep scenarios synthetic and result submissions metadata-only.

Can the same workflow support multiple harnesses?

Yes. Create separate runs with the actual harness and model identifiers for each execution. Keep the plugin version and case set consistent when comparing them.

Start with production onboarding or return to the evaluation test matrix.