
How experiments work
Every experiment has two sides, Baseline (your control) and Experiment (the change you’re testing). You add conditions to each side (see what you can compare on), and Raindrop pulls the matching events and lays the two cohorts next to each other. The results are a set of comparison modules. Each metric shows as a delta bar tagged Better, Worse, or Same, topped by an overall verdict. If the two cohorts are too small or too lopsided to trust, Raindrop says so instead of picking a winner. Experiments compare cohorts of data you’ve already logged; they don’t re-run or replay your agent.Roles and permissions
- Experiments is a Pro feature. On lower plans you can view the page and suggestions as a preview, but creating, editing, and archiving experiments is locked.
- Experiments are scoped to a project. In the all-projects view, creating one prompts you to pick which project it belongs to.
Using experiments
Start from a suggestion
Raindrop refreshes a set of suggested experiments every hour from your recent traffic, so there’s usually a useful comparison waiting for you:- Last 7 days vs. previous 7 days, to catch regressions over time
- Top model vs. all others, or one model head to head with another
- English vs. all other languages, and similar cohort splits
- Explore <signal>, built around your most active negative signals
Build your own
Click New experiment and add conditions to each side. Start with a Date range (the two sides can mirror the same window, split before/after a pivot date, or use independent ranges), then layer on anything else you can compare on: model, tools, signals, custom properties, feature flags, languages, keywords, and more. The empty state walks you through it: choose something to compare and the results fill in.
Read the results
The Summary leads with a verdict (“Experiment is better than Baseline”, “Results are similar”, “Inconclusive, low sample size”) and the top movers between the two cohorts. Below it, modules break down each angle: signal rates, sample size, conversation stats, tool usage, and the model, language, and property mix.
Drill into the events
Click any signal, model, or property row to open a random sample of the matching events on each side, so you can read the conversations behind a delta and confirm it’s real.Save and share
Name an experiment to save it for your team, then Share a scoped link to the comparison. Archive one when it’s no longer relevant; you can restore it later.Reference
What you can compare on
Date range, model, tools, signals, event name, custom properties and user traits, feature flags, input and output keywords, language, and tool errors.What the results show
Signal rates, sample size, conversation stats (length, tools per response, duration), tool usage, and the model, language, and property mix, plus agent-reported self-diagnostics.How Raindrop calls a result
Verdicts are directional, not statistical tests. Raindrop says Inconclusive when a cohort is too small or the two sides differ too much in size, and only calls a winner when the gap is large enough to matter.Troubleshooting
New experiment is locked. Experiments requires a Pro plan. Lower plans can preview the page but can’t create or edit experiments. The verdict says “Inconclusive.” One or both cohorts are too small, or the two sides are too different in size. Widen a date range or loosen the filters so each side has enough events. The two sides aren’t comparable. A “Sample sizes differ significantly” warning means one cohort is much larger than the other. Balance the filters so the cohorts are closer in size before reading the deltas.Related
- Signals to define the behaviors an experiment measures
- Events to explore the traffic behind a cohort
- Triage Agent to spin up a prefilled experiment from a chat investigation