Validation method

Every pre-testing vendor grades their own homework. We seal the exam first.

Similate runs independent validation studies of concept and brand predictors, an AI tool, a vendor, or your own method, against your resolved outcomes and cheap baselines. Predictions are sha256-sealed before outcomes are revealed. We publish no accuracy claims: every number we produce is computed inside a sealed study and belongs to its client and its domain.

1 · Intake, stimuli only

The client supplies historical (or, prospectively, upcoming) tests: test ids, variant ids, and the text of each variant. Outcomes stay with the client. A power estimate states, before anything is sealed, whether the sample can support a decision-grade read; under 20 tests requires a written method-demonstration acknowledgement.

2 · Predictors registered

Any predictor can be audited: an AI tool, a vendor, or the client's own method. Predictions arrive as scores, rankings, or winner-only picks; abstentions are recorded, never hidden. Cheap baselines (longest-text, first-listed, seeded random) are added automatically. One predictor is pre-declared the subject.

3 · Lock, the seal

A sha256 digest is computed over the stimuli, every predictor's predictions, the full scoring configuration, the provenance block, and a pre-registered interpretation rule. The certificate goes to the client BEFORE any outcome is revealed, the client-held copy is the trust anchor, not our database.

4 · Reveal, outcomes enter

Only after the certificate is issued does the outcomes file enter the system. The seal is re-verified (a mismatch voids the run, permanently and disclosed), the optional client outcomes-hash pledge is checked, and a reconciliation table lists every locked test that did not resolve, nothing is silently dropped.

5 · Scorecard & report

The primary metric is separable-pair accuracy, only pairs whose real outcome gap is statistically significant count, because picking among near-ties is noise, not skill. Uncertainty is cluster-aware (bootstrap over tests). Abstentions are never wins. The only verdict language anywhere is the sealed rule's computed verdict, with coverage and sample-grade qualifiers inline.

6 · Right of reply

The audited party may submit one capped, plaintext response after reveal. It is reproduced verbatim in the report, clearly labeled as unverified and outside the sealed run, standard audit practice.

Received a forwarded Similate report? Verify it yourself, trust the math, not us:

  1. Ask the report's named client for the seal certificate they received at lock time (an email with a sha256 digest, timestamp, counts, the sealed interpretation rule and provenance verbatim).
  2. Request the sealed payload JSON for the study from Similate.
  3. Recompute sha256 over the payload's canonical JSON (UTF-8 NFC, recursively sorted keys, no whitespace) with any standard tool, or request our verifier script with the payload.
  4. Compare against the certificate digest. A match proves the predictions, scoring rules, provenance and rule in the report are exactly what existed before outcomes entered the system. A mismatch means the run is void, and voided runs stay on the operator ledger that every report discloses.

Studies are commissioned by insights teams, agencies, or by vendors themselves. Auditee-funded studies are handled the way security audits handle them: a fixed, versioned, non-negotiable protocol; published pricing; no pass guarantee, the report states the result even when the funder loses; and the funding disclosure is printed in the report and is inseverable from any quoted language. The auditee never chooses the test set. Every operator account carries a permanent study ledger (including voided and unrevealed studies) that is disclosed in every report.

What this method is.

A blind, sealed, out-of-sample scorecard of named predictors against resolved outcomes in ONE domain, versus cheap baselines. Predictions, scoring rules, and the interpretation rule are cryptographically sealed before outcomes enter the system; the seal certificate is issued to the client at lock time.

What the seal proves, and doesn't.

The seal proves the predictions in a report match what was sealed before outcomes were revealed to the system. It cannot by itself prove what any party knew beforehand: retrospective studies carry inherent prior-exposure risk (reported as "unknown" unless attested otherwise). Verified prospective locks remove it, predictions are sealed before outcomes exist, verified at reveal against per-test resolution dates. Even then, the seal does not govern how live tests were run or stopped, interim peeking, stopping rules, and which tests resolve remain with whoever runs the tests; the reconciliation table discloses what did not resolve.

What a study is not.

Not causal lift: it scores retrospective ranking of aggregate outcomes; only a live test measures lift. Not a general accuracy claim: a result licenses no claim in any other domain; transfer must be re-proven per domain. Not an in-market launch forecast (behavioral outcomes sit at a specific funnel point) and, for survey-based studies, the ground truth is what respondents SAID (the say-do gap applies to the ground truth itself). Rank-order and calibrated tiers are different claims and are never converted into each other.

Ground-truth caveats.

Outcomes are accepted at face value; their internal validity (randomization, stopping rules, aggregation) is the client's. Historical tests already survived the client's own screening funnel, a restricted stimulus range attenuates measurable skill differences, so a null result does not automatically generalize to the wilder pre-screen concept pool.

When the audited predictor is an AI/synthetic-respondent system.

Independent practitioner guidance (Nielsen Norman Group, "Synthetic Users: If, When, and How to Use AI-Generated 'Research'") warns that AI synthetic respondents are unreliable for concept validation and skew agreeable when reacting to new product concepts. Use front-of-funnel triage plus validation like these studies; do not substitute either for a validated final-launch forecast or live measurement.

Statistical discipline.

Separable pairs within a test are correlated, so primary intervals are bootstrap-over-tests (percentile, 2,000 iterations, seeded, sealed in the scoring config). p-values are per-comparison unless marked Holm-adjusted, with family sizes printed. Small studies are method demonstrations, not decision-grade reads, read the intervals, not the points.

Commission a study

Independent Validation Study: $15,000 fixed, one vertical, one subject predictor, 20–50 resolved tests, seal certificate within 5 business days of complete predictions, report within 15 business days of complete inputs. Vendor bake-off (2–3 predictors, sealed head-to-head): $25,000–$40,000. Prospective locks: same protocol; resolution runs on your test calendar.

winston@similate.ai

Protocol similate-validation-study-v1 · engine validation-studies-v2. Changelog: v2 adds verified prospective locks, external head-to-head with a sealed bake-off rule, exploratory segments, descriptive selective-prediction figures, and the auditee right of reply. The scoring configuration is sealed into every study's digest, a protocol change can never silently apply to an already-locked study.