Measure Your AI Before You Launch It
Evaluate agents, bots, or copilots. Ship what works.
Free
See how your AI performs today
$0/forever
- Bring your own API key: OpenAI, Anthropic, or Together.ai
- One default workspace
- 25 evaluations / month
- 3 context types:
- Scenario (text)
- Conversation (.json up to 512 KB)
- Agent trace / OpenTelemetry (.json up to 1 MB)
- Reference docs (.md only): up to 2 files x 250 KB (500 KB total)
- Single-run evaluations
- Single model per evaluation
- Up to 3 test cases per evaluation
Professional
Most popularReliable scores you can trust
$59/user / month
- Bring your own API key: OpenAI, Anthropic, and Together.ai
- Unlimited workspaces
- Unlimited evaluations (fair use)
- 3 context types:
- Scenario (text)
- Conversation (.json up to 2 MB)
- Agent trace / OpenTelemetry (.json up to 5 MB)
- Reference docs (.md only): up to 5 files x 500 KB (2.5 MB total)
- Single-run and multi-run evaluations
- Statistical multi-run with confidence intervals
- Model comparison (multiple models)
Business
Need higher volume or custom pricing? Talk to us.
Custom
- Everything in Professional
- Higher usage limits
- Custom seat and workspace plans
- Extended data retention
- Admin and support options
- Priority onboarding and training
- Procurement-friendly billing
How It Works
Every plan gives you structured evaluations that validate models quickly, deliver reliable signal, and keep costs predictable.
Fast Validation
Launch evaluations in minutes and compare models side by side.
Reliable Results
Score responses with judge models and confidence intervals.
Multiple Context Types
Evaluate scenarios, conversations, or agent traces — whichever matches your product.
Scores include ±0.3 confidence intervals
Frequently Asked Questions
Everything you need to know about our pricing
An evaluation is one complete run of your prompt against a scenario, conversation, or agent trace — scored by the judge ensemble. On the Free plan you get 25 per month. On Professional, evaluations are unlimited under fair use.
Context types describe what you're evaluating. Scenario is a one-shot text prompt with statistical multi-run so you can see confidence intervals. Conversation lets you upload a multi-turn .json transcript. Agent trace lets you upload an OpenTelemetry .json trace. Reference docs (.md) can be attached to ground the evaluation.
For Conversation and Agent trace, the model output is already provided in the uploaded file. There is no generation path from a primary model. Statistical multi-run applies to Scenario, where Plumloom can regenerate and compare across runs to produce confidence intervals.
Yes. On Free you can bring your own key for OpenAI, Anthropic, or Together.ai. Professional adds the same providers.
Yes. Upgrade instantly and start using new features right away. Downgrades take effect at the start of the next billing cycle.