Reserve
You pick a few decisions you already made and already know the outcome of (headline, subject line, price point, or name A/B). You hold the answer key.
Reserve a few past decisions with the outcomes hidden. We lock a prediction on each , abstaining honestly when it’s too close to call, then you reveal the answer key. You see the hit rate, and the abstentions, before anything changes hands.
The first three founding-partner blind benchmarks are free and in-domain only (media, newsletters, publishers), in exchange for a signed reference, logo, and case-study clause plus participation in a prospective lock. After that, a benchmark is a $15k engagement, refunded if our confident-call accuracy on your data isn't above chance. Pilots and ongoing programs are paid.
Proven in public · frozen, de-leaked benchmark · 1,118 real A/B headline pairs
Point: it gets sharper when it’s sure, and it can pass when it isn’t. GPT-5 stays near 57% either way. This is public media click-ranking, not a promise about your funnel. The only number that counts is the hit rate on your decisions, locked before reveal.
How it works
You pick a few decisions you already made and already know the outcome of (headline, subject line, price point, or name A/B). You hold the answer key.
We rank each option with a learned model and seal the prediction with a timestamped hash before any outcome is shared. It can return “too close to call.”
You reveal the real winners. The seal proves the prediction could not have been changed after the fact.
We score it together: where we picked the winner, where we abstained, where we missed, with confidence intervals and nothing hidden.
Why it survives diligence
Every prediction is sealed with a cryptographic hash and a timestamp before you reveal a single outcome. Tamper with it and the seal fails. You never have to take “we didn’t fit it after the fact” on faith.
When two options are not statistically separable, the model says so instead of guessing. Abstentions are reported in full. They are what makes the hit rate trustworthy.
A learned ranker over transparent copy features. You can see which phrases moved each score. Not an opaque survey panel, not a large language model improvising.
The model is validated on public media-copy click outcomes. Whether that transfers to your domain and your metric is exactly what this benchmark measures, on your decisions, before you pay.
What the public numbers mean
On 1,118 public headline A/Bs, Similate climbs when the call is clearer. GPT-5 stays near 57%.
Same pairs · Similate vs GPT-5
Scope: engagement ranking on public media headlines, not causal conversion lift (decision-grade n=192). Full methodology on request.
Start with one decision
Bring a campaign, pricing, positioning, or naming call you already made. If we don’t beat your gut, you’ve spent twenty minutes. If we do, we talk about a paid pilot on a live decision.