> ## Documentation Index
> Fetch the complete documentation index at: https://raindrop.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Test your agent from Raindrop

> Set up your harness, create an eval, configure a replay suite, run it, and improve your agent in the playground.

A custom replay endpoint lets you test **your real agent** from Raindrop. Pick an input dataset, choose the evaluators, and run. Your server generates the replies; Raindrop shows the outputs and grades them.

You keep your agent code, model credentials, and tools in your own application. The SDK provides a small wrapper around the function you already use to call your agent.

Follow this flow:

1. **Set up your harness** — connect your existing agent to Raindrop.
2. **Create an eval** — describe what a good reply looks like.
3. **Configure your replay suite** — choose the agent, inputs, and evaluators.
4. **Run** — generate replies and review the scores.
5. **Use the playground** — try changes on a few rows, then run the full suite.

You usually set up the harness once. After that, most of the work happens in Raindrop.

## 1. Set up your harness

A **harness** is a small adapter around your agent that lets Raindrop call it with test inputs. Your agent still runs on your server.

### What the harness does

One URL, such as `https://your-app.com/api/evals`, handles two jobs:

* **Discover your agents:** Raindrop sends a `GET` request to learn which agents and parameters to show in the form. This does not run your agent.
* **Run your agent:** Raindrop sends a `POST` request when you start a test. The SDK loads the selected dataset rows, calls your agent for each row, and sends the traces back to Raindrop.

You do not need to build a dataset loop, manually send row IDs, or implement the discovery response yourself.

### Install the SDK and prepare your keys

```bash theme={null}
npm install raindrop-ai zod@3
```

Use an SDK release that exports `defineAgent` and `createEvalHandler`. If your installed version does not include them, use the replay-enabled release provided by your Raindrop team before continuing.

Add these variables to your server environment:

```bash theme={null}
RAINDROP_API_KEY=your-raindrop-organization-api-key
RAINDROP_PROJECT_ID=default
EVAL_ENDPOINT_SECRET=a-long-random-secret-you-create
```

These credentials do different jobs:

* **`RAINDROP_API_KEY`** lets the SDK read the run and upload its traces to Raindrop. Use an [organization API key](https://auth.raindrop.ai/org/api_keys), not a telemetry write key. The key must have access to the project you are testing.
* **`EVAL_ENDPOINT_SECRET`** protects your endpoint. You will enter this same secret in Raindrop's endpoint configuration so Raindrop can call your server.

Your agent's model-provider keys stay configured as they are today. None of these secrets belongs in browser code or a dataset.

### Wrap your agent

Here is a complete **Next.js App Router route**. Put it in `app/api/evals/route.ts`, then replace the `answerCustomer` import and call with your own agent function.

The example assumes your agent accepts a message and options, then returns its final reply as a string.

```typescript theme={null}
import { Raindrop, createEvalHandler, defineAgent } from "raindrop-ai";
import { z } from "zod";

import { answerCustomer } from "@/lib/agent";

export const runtime = "nodejs";

function requiredEnv(name: string): string {
  const value = process.env[name];
  if (!value) throw new Error(`Missing ${name}`);
  return value;
}

const endpointSecret = requiredEnv("EVAL_ENDPOINT_SECRET");
const raindrop = new Raindrop({
  apiKey: requiredEnv("RAINDROP_API_KEY"),
  projectId: requiredEnv("RAINDROP_PROJECT_ID"),
});

const support = defineAgent({
  slug: "support",
  name: "Support agent",
  description: "Answers customer support questions.",
  parameters: {
    prompt: z.string().default("Be helpful, clear, and concise."),
    temperature: z.number().min(0).max(2).default(0),
  },
  run: async (row, { parameters }) => {
    if (row.input === null) {
      throw new Error(`Dataset row ${row.id} has no input`);
    }
    const input = row.input;

    return raindrop.tracer({ user_id: "eval-runner" }).withSpan(
      { name: "support", inputParameters: [input] },
      async () => {
        const reply = await answerCustomer(input, {
          systemPrompt: parameters.prompt,
          temperature: parameters.temperature,
        });
        return reply;
      },
    );
  },
});

const handleEval = createEvalHandler(raindrop, { agents: [support] });

async function handleRequest(request: Request): Promise<Response> {
  if (request.headers.get("authorization") !== `Bearer ${endpointSecret}`) {
    return new Response("Unauthorized", { status: 401 });
  }
  return handleEval(request);
}

export const GET = handleRequest;
export const POST = handleRequest;
```

### What each wrapper does

**`defineAgent` describes the agent.** Its `name` appears in the Agent dropdown. Its `slug` is the stable identifier used to call it. Its `run` function receives one dataset row at a time.

**`withSpan` records your agent call.** It captures the input and the returned reply. Keep the agent's work inside the callback and await it. If your agent returns a response object, return the actual final text from that object. If it streams, collect and return the complete reply.

If your agent already uses Raindrop tracing, keep that instrumentation. You do not need to add a second copy of its existing root wrapper; the replay handler associates work performed inside `run` with the right dataset row. Existing model and tool instrumentation adds detail to the trace.

**`createEvalHandler` supplies the HTTP handler.** It handles discovery, validates parameters, claims the run, runs the rows, and uploads the results. Keep both `GET` and `POST` behind the same authentication.

Using a different framework? Mount `handleRequest` on a route that accepts a standard Web `Request` and returns a `Response`, or use your framework's adapter. The agent definition stays the same.

### Check agent discovery

Start your application and call the route:

```bash theme={null}
curl --fail-with-body http://localhost:3000/api/evals \
  -H "Authorization: Bearer $EVAL_ENDPOINT_SECRET"
```

You should receive JSON containing an `agents` array with your `support` agent, its name, and the `prompt` and `temperature` fields. A current handler also returns `concurrencySupported: true`.

This is the same discovery request Raindrop makes. If this check fails, fix it before starting a test:

* **401:** the endpoint secret is missing or does not match.
* **404 or HTML:** the URL is not pointing to the replay route.
* **An empty agent list:** check the `agents` array passed to `createEvalHandler`.

Do not invent a run ID to test `POST`. Start a test in Raindrop; it creates the run before calling your endpoint.

Your harness is ready when discovery returns your agent and its parameters. Deploy it to a server reachable by Raindrop over HTTPS. For local development, use an authenticated HTTPS tunnel; hosted Raindrop cannot reach `localhost` on your laptop.

## 2. Create an eval

An **eval** checks one aspect of your agent's reply. Start with a single clear question, such as “Did the agent answer the user's question without unnecessary detail?”

In **Evals**, create an evaluator and describe what it should check in plain language. For example:

> Check whether the reply is concise and helpful. It should answer the user's question directly, avoid repeating itself, and include only the detail needed to take the next step.

Review the generated policy and adjust it to match your expectations. If you have examples or human-labelled comparisons, attach them with **Add reference**. For a judge built from human labels, use **Calibrate** to check its agreement with those labels before relying on its scores.

Choose the kind of check that matches your goal:

* **Score one reply:** assess helpfulness, correctness, or another quality independently.
* **Compare two replies:** use a pairwise evaluator to decide whether a new reply is better than a reference.

You are ready for the next step when your evaluator is created and available to select. You can add more later—for example, a coaching preference judge alongside a word-count eval.

## 3. Configure your replay suite

A **replay suite** brings together your agent, an input dataset, and the evals that will check its replies. A **run** is one execution of that setup.

If you do not have an input dataset yet, create or upload one in **Datasets**. Start with around 10 representative inputs. Inputs alone are enough to generate replies; you do not need to write example answers just to get started.

Open **Evals** and the **Test your agent** form.

1. Keep **Run in → Custom endpoint** selected.
2. Under **Agent endpoint**, choose **New endpoint**.
3. Give it a recognizable name, such as **Support staging**.
4. Enter the full route URL: `https://your-app.com/api/evals`.
5. Add a header with the name **`Authorization`** and the value **`Bearer YOUR_ENDPOINT_SECRET`**. Replace `YOUR_ENDPOINT_SECRET` with the value of `EVAL_ENDPOINT_SECRET` on your server. Include `Bearer` and the space.
6. If the **Concurrency** field is available, set it to the number of dataset rows your harness can handle at once. Save the endpoint.
7. Choose **Support agent** from the discovered agents.
8. Choose your **Input dataset**. Start with around 10 rows so you can check the setup quickly.
9. Expand **Parameters** if you want to change the prompt or temperature.
10. Select the **Evals** you created and check the name. It starts with your dataset name, and you can edit it.

The Authorization header uses your **endpoint secret**. Your **Raindrop API key** stays on your server, where the SDK uses it to read the run and upload traces.

If your suite includes a pairwise eval, it needs something to compare each new reply against. Raindrop uses the dataset's reference when available. Otherwise, the form explains that the first run will create a baseline for the next run. For a labelled A/B dataset, the expected winner supplies the reference, with Answer B used for ties.

You are ready to run when the agent, input dataset, and evals are selected.

## 4. Run

Click **Run suite**. You will land on the run, where you can follow each input through execution and grading.

Raindrop fixes the dataset version and run configuration, then calls your endpoint. The SDK runs your agent on those inputs and uploads the traces. The selected evaluators grade the resulting outputs separately.

**Execution completed** means your agent finished. Grading may still be running. Your endpoint does not need to call the evaluators itself.

Concurrency controls how many **dataset rows your agent handles at once**, not how many judges run. Start low if your agent uses shared state or has model rate limits.

Open individual results to inspect the input, generated reply, and evaluator outcome. Check a few examples yourself before drawing conclusions from the overall scores.

When grading finishes, you have a recorded run to return to and compare with future runs.

## 5. Use the playground

Open the suite's **Playground** when you want to try a change quickly.

1. Select a few dataset rows, or enter a sample percentage and click **Sample**.
2. Expand **Parameters** and change the prompt or another setting your harness exposes.
3. For a pairwise comparison, select a baseline run when you want to compare against its saved outputs. Without a selected baseline or a dataset reference, the pairwise cells show **No baseline selected**; you can still generate replies and run the other evals.
4. Click **Run selected** and review the outputs and scores for those rows.
5. Adjust and repeat. Use the same rows when comparing two prompt changes so you can see what actually improved.
6. When you are happy with the change, click **Run suite** to run the full input dataset with your current configuration and record it in suite history.

You do not need to save the parameters as a preset and find them again to move from the playground to a full run. **Save** is there when you want a named configuration to reuse later.

## Optional harness features

The five steps above are the core flow. These options help when you want to expose more controls or support more agents.

### Add more parameters

The `parameters` declaration is also the form definition: Raindrop discovers it and renders the controls automatically.

You can declare text, number, boolean, and choice fields with `z.string()`, `z.number()`, `z.boolean()`, and `z.enum()`. Use `.default()` for sensible starting values and `.describe()` for a short explanation. These examples use Zod 3.

Parameter values arrive in `run` as `parameters`. They only affect your agent if you pass them into your agent function, as the example does with `systemPrompt` and `temperature`.

You can edit values for a test without saving a reusable preset. Use **Save** when you want to reuse a named configuration later. This makes it easy to try one prompt, change it, and run the same inputs again.

### Register more than one agent

Register them on the same endpoint:

```typescript theme={null}
const handleEval = createEvalHandler(raindrop, {
  agents: [support, sales, onboarding],
});
```

Each entry is a `defineAgent(...)` result with a unique slug. Raindrop discovers all of them and shows their names in the Agent dropdown. Each agent can expose different parameters.

### Share setup across rows

Use `setup` and `cleanup` when a test needs a sandbox, temporary workspace, or other shared resource:

```typescript theme={null}
const support = defineAgent({
  slug: "support",
  parameters: { prompt: z.string().default("Be concise.") },
  setup: async () => createTestWorkspace(),
  run: async (row, { parameters, environment }) => {
    return runTracedSupportAgent(row.input, {
      prompt: parameters.prompt,
      workspace: environment,
    });
  },
  cleanup: async (environment) => environment.dispose(),
});
```

Here, `createTestWorkspace` and `runTracedSupportAgent` stand for your application's setup and traced agent functions. Setup happens once when execution reaches the first row. Rows share the returned environment, and cleanup runs after they settle, including after row failures. Keep each row's conversation state separate. Cleanup cannot run after a process is killed, so temporary resources should also expire on their own.

## If something does not work

### Discovery works, but starting the run returns 401

Check which request failed. A 401 from **your route** means its Authorization header is wrong. A credential error when the **SDK calls Raindrop** means `RAINDROP_API_KEY` is invalid, expired, or cannot access the selected project. A successful discovery request does not verify the SDK's Raindrop credentials.

### The run has no usable output

Return the final reply from your agent callback and await all the work that produces it. Confirm tracing is enabled and your traced call includes the input and output. Returning early while work continues in the background can leave the run without a complete trace.

### The request times out

The handler keeps the request open until agent execution finishes. Run it on infrastructure whose request timeout can cover the whole batch, including any proxy in front of it. A short-lived serverless route may be unsuitable for larger datasets. Do not return an early success response or put `handleEval` into an unawaited background task.

Start with fewer rows, check the slowest agent call, and adjust concurrency within your agent's limits. Raindrop's current endpoint execution request timeout is 15 minutes.

### A parameter is rejected

Check its declared type and limits. For example, `temperature` must be a number within the schema's range. The SDK rejects invalid parameters before running the agent. Custom parameter values must be strings, finite numbers, or booleans; objects, arrays, and `null` are not supported.

### The agent list does not show a new agent

Deploy the code that registers the agent, then check `GET` on the exact endpoint URL saved in Raindrop. If the returned list is correct, select that endpoint again to refresh discovery.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.