Why no percentages · Windrose Atlas
An honest instrument

Any AI can print "your conversion will be 23%." It will be wrong.

The category sells false precision. We sell a direction you can trust, and a label that tells you exactly how far to trust it. This is the argument, in full.

Simulated research has a dirty secret: the numbers look precise and aren't. Peer-reviewed work shows that language models do not reliably reproduce human psychology, absolute predictions drift, distributions collapse, and confident decimals hide guesswork. We built Windrose Atlas on that finding, not against it.

So here is our deal with you. Without real-world data, Windrose Atlas reports three things it can defend: the direction of an effect, the ranking of your variants, and the hypotheses most worth your scarce time.

Every number ships with an evidence label, what rung it stands on, whether it's calibrated, and where our backtests say the simulator can be trusted. When we don't know, the product says "we don't know." That's not a disclaimer. That's the product.

Bring your own outcomes later, sales, sign-ups, past experiments, and the label climbs from SIM-RANK to a CALIBRATED rung. Only your real data earns that, never us.

The obvious question

Why not just ask Claude? Three structural reasons, and a confession.

1

The optimisation target points the wrong way. Frontier assistants are trained to be helpful, rational and agreeable. That training systematically removes the impulsiveness, bias, inconsistency and variance that make real customers real. A model rewarded for not being irrational cannot simulate an irrational buyer, and a model rewarded for agreeableness will tend to like your idea. Our oracle is optimised in the opposite direction: a cognition model trained on millions of documented human decisions from published behavioural experiments, deliberately kept away from assistant training. This isn't a gap a better prompt can close. It's a different objective function.

2

An answer is not a measurement. A chat gives you one sample from a model sensitive to phrasing, ordering and temperature, with no experimental design behind it. A Windrose run is a designed experiment: a factorial grid of personas and conditions, counterbalanced wordings, dispersion checks across paraphrases, one hundred resamples behind every ranking, and a stability score instead of a vibe. Same input, same instrument, same artefacts, every run auditable to a manifest hash. Conversations are not reproducible. Instruments are.

3

A chat never finds out it was wrong. There is no loop from an assistant's guess back to what the market actually did. Windrose closes that loop: clients and agents return real outcomes, each becomes a paired (simulation, reality) observation, and calibration runs from real to sim, never the reverse. That growing corpus of verified pairs exists inside no foundation model and cannot be scraped. It is the asset.

The confession: we use Claude. A frontier model designs our scenarios and sits in the ensemble as a second, independent oracle; the agreement between oracles is reported on the label. We don't compete with frontier intelligence, we build the instrument around it. When the models get better, the instrument gets better. What they will never grow on their own is the method, the epistemic contract enforced in code, and the verification corpus.

A brilliant consultant gives you an opinion. A laboratory runs the experiment that can prove you wrong. Windrose is the laboratory.

What we report instead

Three things we can defend under scrutiny.

DIRECTION

Which way the effect points.

Does this message help or hurt? Directional, and honest about its scope.

RANKING

Which variant is stronger, for whom.

Order over decimals, with a stability score for how firmly it holds.

HYPOTHESES

What to test in the real world next.

The bridge to fieldwork, the conversations worth having first.

When we don't know, the product says so. That's not a disclaimer, that's the product.

The methodical questions.

How accurate is this?

We report direction and ranking, not absolute numbers. Every result carries a backtest class and an evidence label showing exactly where the simulator can be trusted, and where it can't.

Why should I trust a simulation?

You shouldn't, blindly. That's why every result carries its rung on the evidence ladder. A simulation selects which hypotheses to test first; it never stands in for real-world data.

Is this a replacement for talking to customers?

No. It tells you which conversations to have first.

What data is the population built on?

Census demographics and published survey attitudes, parameterised by documented cultural frames. The full source list lives on the Science page.

See the whole method, then run your own.

The three axes of uncertainty, the backtest protocol, the full label, all on the Science page.