How MarketingBench Works
An independent, open benchmark for AI marketing quality. The protocol, the math, the safeguards, and the limitations, all on one page.
1. Blind evaluation
Two anonymous AI models receive the same marketing task with the same system prompt. Voters see only “Response A” and “Response B”, with no model names or provider branding. You pick the better output. Only after the vote is cast are the models revealed.
Same prompt. Same instructions. No labels. Pure output quality.
2. How ratings work
We use Elo ratings, the system used in chess. Every model starts at 1200. Winning a matchup gains points, losing costs points, and the amount depends on the opponent: beating a top-ranked model is worth more than beating a low-ranked one (K-factor 32).
Every rating is published with a 95% confidence interval that narrows as votes accumulate. Models with fewer than 30 votes are explicitly flagged provisional. When two intervals overlap, the models are statistically tied. Read them that way.
3. What we test
General AI benchmarks test code, math, and chat. They say very little about which model writes the best subject line or the most on-brand push notification. MarketingBench tasks mirror real production work, with the format constraints those channels actually impose:
- Copywriting — headlines, taglines, product copy, landing page sections
- Search ads — Google RSA-style headlines and descriptions within character limits
- Push notifications — short-form hooks with strict length constraints
- SMS — promotional and lifecycle texts with compliance-aware formatting
- Translation — faithful and transcreation modes for marketing copy
Every model gets the exact same system prompt for each task type. All of them are public on the System Prompts page.
4. Keeping votes honest
- Blind until voted. Model identities are only revealed after your vote is recorded, so there is no way to vote for a brand.
- One vote per matchup. Duplicate votes on the same matchup are rejected.
- Fingerprint de-duplication. Anonymous votes are de-duplicated using a hashed (SHA-256) fingerprint; raw IPs are never stored against votes.
- Rate limits. Voting is rate-limited per account and IP to blunt scripted voting.
- Small samples flagged. Provisional badges and minimum-vote thresholds mark ratings that do not yet have enough votes to mean much.
5. Full transparency, no pay-to-play
Vote counts, win rates, confidence intervals, and head-to-head records are all public on the Leaderboard. There is no paid placement, no editorial weighting, and no sponsored ranking. MarketingBench is independent: no affiliation with any AI lab or marketing software vendor. When a model ranks well here, it earned it blind.
Tool recommendations: salience is not product quality
The Tool Finder asks one current model from each provider in the pinned benchmark panel to rank three substitutable platforms for the same marketing subjob and company constraints. Provider-balanced results prevent several models from one lab from looking like independent agreement.
Open-ended scans measure which products a model recalls and suggests without seeing a candidate list. Controlled scans randomize a supplied shortlist and measure preference only within that set. Those results are reported separately and are never described as an objective “best tool” ranking.
Models do not verify current pricing, integrations, security claims, or feature availability during a scan. Treat the output as a structured research starting point, then confirm important claims directly with vendors before buying.
Prompt trends: convergence is not verified quality
Prompt Trends asks one current model from each provider in the pinned benchmark panel which reusable prompts a marketer should keep on hand for a standardized marketing method and tool context. Each model ranks three prompt archetypes from a fixed, versioned catalog and writes its best verbatim prompt for each, with archetype order randomized per run.
Only runs where the full provider-balanced panel returned valid structured output are published, and results are aggregated by archetype so convergence across labs is measurable. The verbatim prompts shown are exactly what each model returned — they are model output, not editorially reviewed templates.
A prompt many models recommend is a strong starting point, not a guaranteed performer. Replace the placeholders with real inputs, test the output against your own quality bar, and keep what earns its place.
6. Limitations: read this before citing the rankings
- Preference is not performance. Blind votes measure what marketers prefer, not click-through or conversion rates. Preference is a fast, useful quality signal, but the two can diverge. Evaluation against real campaign outcomes is planned.
- Sample sizes matter. A ranking with wide confidence intervals is a hypothesis, not a verdict. We show the intervals so you can judge for yourself.
- Models move. Ratings reflect the current API versions of each model; providers ship updates and rankings shift. The rating history chart makes those shifts visible.
- Coverage is copy-first. Today the benchmark covers short-form marketing writing. Strategy, analytics, and agentic campaign execution are separate tracks in development.
Frequently asked questions
- How is MarketingBench different from LMArena (Chatbot Arena)?
- LMArena measures general chat preference across every kind of prompt. MarketingBench evaluates one thing: marketing work product (ad copy, email, push notifications, SMS) under the format constraints those channels impose. It is the vertical complement to general arenas, not a competitor.
- Can model providers pay to improve their ranking?
- No. There is no paid placement, no sponsored ranking, and no way to buy votes or visibility. Rankings are computed only from blind pairwise votes. MarketingBench is independent and is not affiliated with any AI lab or marketing software vendor.
- How do you prevent vote manipulation?
- Model names are hidden until after a vote is cast, each matchup accepts one vote per voter, anonymous votes are de-duplicated with a hashed fingerprint, and voting is rate-limited per user and IP. Models with fewer than 30 votes are flagged as provisional, so a lucky streak on a handful of votes does not read as a ranking.
- How many votes does a ranking need before it is reliable?
- Every rating is shown with a 95% confidence interval that narrows as votes accumulate. We flag models with fewer than 30 votes as provisional, and pairwise win rates in the head-to-head matrix are only displayed once a pair has 10+ votes. When intervals overlap, treat the models as statistically tied.
- Do votes predict real-world marketing performance?
- Not necessarily. Votes measure what marketers prefer in a blind comparison, which is not the same as click-through or conversion rate. Preference is a fast, useful quality signal, but the two can differ. Evaluation against real campaign metrics is planned.
- Can I test my own prompts?
- Yes. The Playground runs your own marketing prompt through anonymous models, and the Evaluation Builder adds reusable weighted criteria. Custom prompts, rubric scores, and decisions are private and unranked, so they never skew the public leaderboard.