Field NotesSEO

The Elegant Mathematization of AEO

By October 8, 2026No Comments11 min read

I was reading James Webb Young’s How to Become an Advertising Man, published in 1963, and stumbled on a word I didn’t expect: scientism.

Young was an advertising executive. He was the first chairman of The Advertising Council, had spent decades at J. Walter Thompson, and wrote the book as a practical guide to the profession.

Buried in the practical career guide was a lesson in epistemology.

I’d last encountered the term in Nassim Taleb, who uses it frequently and with characteristic venom.

Finding the same word in a 60 year old advertising manual suggested the problem was older and more pervasive than I’d assumed.

Young was warning his industry about something that Taleb would later warn everyone about. And then I went down the rabbit hole and realized it is deeper than either of these sources.

The Slavish Imitation

The term “scientism” was popularized by Friedrich Hayek.

In a series of essays published in Economica during World War 2, later collected as The Counter-Revolution of Science in 1952, Hayek defined scientism as “the slavish imitation of the method and language of Science.”

Basically, it’s the method and language of science, applied to domains where those methods do not produce useful knowledge.

Hayek’s target was social planning: the belief, common among mid-century technocrats, that if you could just measure society with sufficient precision, you could engineer it rationally. Basically, the rigorous methods of physics and chemistry, transplanted to economics and governance.

His argument wasn’t against measurement, but against the misapplication of it. The social world is too complex, too reflexive, too human for the methods of the laboratory to produce the certainty they produce in the laboratory.

Before Hayek, Auguste Comte’s positivism had laid the groundwork for the idea that all genuine knowledge is scientific knowledge, that any question not amenable to scientific method is not a real question.

After Hayek, A.J. Ayer’s logical positivism pushed the idea further: any statement that cannot be empirically verified is literally meaningless.

There must be some nuance here, right?

The Costume of Rigor

“Some fields of research have a particular culture about how work is expressed or assessed, which can further amplify problems. In academic economics, for example, it is common to present even simple concepts as detailed equations. Ecologist Robert May once called it the ’elegant mathematization of things,’ giving ideas an appearance of rigor that they may not deserve.”

— Adam Kucharski, Proof: The Art and Science of Certainty

Adam Kucharski’s book Proof is about certainty, as well as how humans establish truth, gather and weigh evidence, and navigate uncertainty across mathematics, science, law, politics, and artificial intelligence.

I’ve said it before and I’ll say it again: All decisions include some level of uncertainty. Properly collected and analyzed data can reduce that uncertainty, but never eliminate it.

So, this book covers that grey area that all of us, especially marketing leaders operating within AI search, deal with, where you are neither 100% or 0% certain, but somewhere in the middle.

The problem many are dealing with is that there is genuine uncertainty in how to measure, influence, and approach organic growth inside AI engines. And then there is an abundance of studies, many of which go viral on LinkedIn, than purport to have clean, crisp numbers to back up strategic takeaways (which are limitless – Reddit, LinkedIn creator programs, syndication, press releases, comparison pages, high scale listicle publication, no listicles at all, and on and on).

It’s confusing. 

Here are a few numbers I’ve seen, for example:  

Domain Rating has a Spearman correlation of 0.32 for AI visibility, and branded search volume has a Spearman correlation of 0.39. Comparison content has a Spearman correlation of 0.56 for AI referral traffic.

All of this sounds very rigorous, but is it meaningful?

Let me just riff a bit here on the limitations with these hard numbers:

  • Sampling is hugely consequential. What sites? What industries? B2B vs. consumer? Famous brands vs. unknown brands? English only? Which prompts? If you change the prompt universe, you may change the result substantially. There is no canonical population of “AI searches” from which researchers can randomly sample.
  • AI visibility itself is a noisy measurement. Visibility depends on the prompts researchers choose, model, model version, geography, personalization/context, date, repeated runs, and how “visibility” is scored. Researchers aren’t directly observing market wide AI exposure. They’re constructing an index from a sample of synthetic queries.
  • The independent variables are proxies too. Domain Rating isn’t “authority”; it’s Ahrefs’ proprietary link metric, itself a correlation of sorts. Branded search volume isn’t brand strength; it’s one imperfect manifestation of it. So you’re correlating one proxy with another proxy.
  • Confounding is everywhere. High DR websites also tend to be older, better known, better funded, more frequently mentioned, covered by journalists, cited on Reddit, searched more often, and have vastly more content. Which of those actually causes AI visibility? A bivariate Spearman correlation can’t tell you.

So this is where Robert May’s observation about academic economics comes in. The equations are not wrong, and the math checks out. But the mathematization gives the idea an appearance of precision that the underlying phenomenon does not support.

These aggregate studies are useful in that they can reveal patterns, generate hypotheses, and tell us where to look. The mistake is asking them to answer questions their design cannot answer.

The magnitude of the correlation is almost beside the point. The research design often cannot support the strategic interpretation being placed on it.

A model can be numerically precise while its causal interpretation remains profoundly uncertain.

The Marketing Version

I wrote a few months ago about patterns in the clouds, which covered the problem of cherry picking, spurious correlations, and industrial scale data mining in AI search research.

Basically, when you have a massive amount of data and very little alignment on what to measure, you can prove nearly anything.

It’s sort of the Texas Sharpshooter fallacy, where a cowboy shoots a gun randomly at the side of the barn, then walks up to the wall and paints a bullseye around the tightest cluster of bullet holes, claiming he is an absolute expert marksman.

Image Source

This essay is about a related but different problem. Not bad data, but the wrong epistemology. The belief that the right response to a complex, non-deterministic, rapidly shifting system is always more measurement, more precision, more analysis. That if we can just build a large enough dataset of prompts, track enough citation sources, run enough regressions, we will converge on the “correct” strategy for AI search with something approaching scientific certainty.

I wrote about this in Probability Engineering, the idea that LLMs are probabilistic systems, not deterministic ones, and that the appropriate response to a probabilistic system is not to pursue certainty but to manage distributions.

You Can Measure Anything (…Should You?)

“Your scientists were so preoccupied with whether they could, they didn’t stop to think if they should.” ― Dr. Ian Malcolm, Jurassic Park

Measurement is awesome. Huge fan. Y’all, I’ve spent much of my career reading statistics textbooks, building models in R and Python, and applying the knowledge through controlled experiments.

This isn’t an argument against measurement. It’s about the misapplication of measurement to make something murky seem certain through mathematical language.

It is an argument for knowing what kind of problem you are facing before you choose your instrument.

Some problems are amenable to scientific method in the strict sense:

  • Testing a landing page headline.
  • Measuring the effect of page speed on conversion rate.
  • Comparing two email subject lines.

These are bounded systems with controllable inputs, measurable outputs, and repeatable experiments. The A/B test is the right tool.

Some problems respond to pattern recognition:

  • Tracking how citation sources shift across AI platforms over a time series.
  • Monitoring which competitors appear in model responses for your category.
  • Watching how your brand’s share of voice changes quarter to quarter.

The data here is useful as a compass (directional, approximate, worth watching) but not as a coordinate. The aggregate tells you something, but not what specifically to do without triangulation, mechanistic understanding, and quite frankly, first principles reasoning.

Some problems require judgment (calibrated by data):

  • Choosing which categories to build authority in.
  • Deciding how to position your brand in a market that AI models are still learning to describe.
  • Predicting which content formats will matter in a year.

These are decisions where the quality of your reasoning and the depth of your experience matter more than the size of your dataset. The data can inform, but it can’t determine. The person who has spent years in the category, who understands the buyers, who has taste for what resonates will likely make a better decision with a small dataset than a newcomer will make with a large one. Here, I suggest brushing up on bayesian reasoning but also having conviction.

And in some areas, measurement may fail to isolate the causal inputs or may compress a multidimensional construct into a misleading scalar. I’m sure you could measure love, but perhaps it’s worth keeping it a magical mystery. Brand trust is such a composite outcome that isolating inputs is likely to cause more noise than signal. Does employee NPS measure the quality of a culture? I have my doubts.

The point is, you should have many tools and apply the right instrument to the right problem.

If you want to understand how a page might be made more retrievable for a defined set of prompts, you can use retrieval-oriented testing: passage relevance, embeddings, lexical matching, reranker evaluation, and repeated tests against the systems where possible.

If you want to know why Nike SB wins visibility for “best skate shoe,” these tools may not be as useful. Different questions require different instruments.

The Advertising Man

James Webb Young spent a career in a field that was becoming increasingly data-driven and he understood the value of that shift. But he also understood that advertising (like marketing, like brand building, like navigating a system as complex and unpredictable as AI search) is a practice that requires judgment.

Judgment is not the absence of data.

Judgement is the thing that decides what the data means, what to do with it, and when to set it aside.

Hayek’s critique was that the language of science could be used to shut down legitimate forms of knowledge such as local knowledge, tacit knowledge, the kind of understanding that practitioners develop through experience and cannot easily articulate in a paper.

The version for our industry is fairly apparent if you spend much time on LinkedIn (which I do often warn about).

The data is proliferating, and the tools are improving, and the temptation is to treat all AI search problems as problems wherein Spearmans and aggregate citation shifts can provide meaningful strategic insight.

Know when to measure. Know when to observe. Know when to reason. Know when to trust your judgment over someone else’s dataset.

Want more insights like this? Subscribe to Field Notes

Alex Birkett

Alex is a co-founder of Omniscient Digital. He loves experimentation, building things, and adventurous sports (scuba diving, skiing, and jiu jitsu primarily). He lives in New York City with his dog Biscuit.