Illustrative campaign example · not currently accepting applicants
AI evaluation
PromptBench
An evaluation workspace for comparing prompts, models, datasets, costs, and quality regressions across every AI release.
$38 per approved outcomeActivated account
PromptBench workspace
Pass rate
94.2%
Helpfulness01
Factuality02
Schema match03
Creator brief
What a strong creator would make
Build a small evaluation suite for a support copilot or extraction workflow. Demonstrate a regression, compare models, and show how a team decides what is safe to release.
Best audience: AI engineers, product teams, and founders shipping LLM features