ResearchBlogHelp & SupportLearn EvalsComing Soon
    Pricing

    Measure Your AI Before You Launch It

    Evaluate agents, bots, or copilots. Ship what works.

    MonthlyYearly20% Off

    Free

    See how your AI performs today

    $0/forever

    • Bring your own API key: OpenAI, Anthropic, or Together.ai
    • One default workspace
    • 25 evaluations / month
    • 3 context types:
      • Scenario (text)
      • Conversation (.json up to 512 KB)
      • Agent trace / OpenTelemetry (.json up to 1 MB)
    • Reference docs (.md only): up to 2 files x 250 KB (500 KB total)
    • Single-run evaluations
    • Single model per evaluation
    • Up to 3 test cases per evaluation

    Professional

    Most popular

    Reliable scores you can trust

    $59/user / month

    • Bring your own API key: OpenAI, Anthropic, and Together.ai
    • Unlimited workspaces
    • Unlimited evaluations (fair use)
    • 3 context types:
      • Scenario (text)
      • Conversation (.json up to 2 MB)
      • Agent trace / OpenTelemetry (.json up to 5 MB)
    • Reference docs (.md only): up to 5 files x 500 KB (2.5 MB total)
    • Single-run and multi-run evaluations
    • Statistical multi-run with confidence intervals
    • Model comparison (multiple models)

    Business

    Need higher volume or custom pricing? Talk to us.

    Custom

    • Everything in Professional
    • Higher usage limits
    • Custom seat and workspace plans
    • Extended data retention
    • Admin and support options
    • Priority onboarding and training
    • Procurement-friendly billing

    How It Works

    Every plan gives you structured evaluations that validate models quickly, deliver reliable signal, and keep costs predictable.

    Fast Validation

    Launch evaluations in minutes and compare models side by side.

    Reliable Results

    Score responses with judge models and confidence intervals.

    Multiple Context Types

    Evaluate scenarios, conversations, or agent traces — whichever matches your product.

    Evaluation Results
    Model ComparisonScenario · multi-run
    A
    gpt-5.5
    4.6
    B
    claude-opus-4.7
    4.4
    C
    gemini-3.5-flash
    3.8

    Scores include ±0.3 confidence intervals

    Frequently Asked Questions

    Everything you need to know about our pricing

    An evaluation is one complete run of your prompt against a scenario, conversation, or agent trace — scored by the judge ensemble. On the Free plan you get 25 per month. On Professional, evaluations are unlimited under fair use.

    Context types describe what you're evaluating. Scenario is a one-shot text prompt with statistical multi-run so you can see confidence intervals. Conversation lets you upload a multi-turn .json transcript. Agent trace lets you upload an OpenTelemetry .json trace. Reference docs (.md) can be attached to ground the evaluation.

    For Conversation and Agent trace, the model output is already provided in the uploaded file. There is no generation path from a primary model. Statistical multi-run applies to Scenario, where Plumloom can regenerate and compare across runs to produce confidence intervals.

    Yes. On Free you can bring your own key for OpenAI, Anthropic, or Together.ai. Professional adds the same providers.

    Yes. Upgrade instantly and start using new features right away. Downgrades take effect at the start of the next billing cycle.