We Were in the Room When the Numbers Got Made
Almost every Big 4 and MBB report you have ever cited was built backwards. We know, because we helped build them.
One of our contributors once caught a polling company fabricating respondents.
The firm had commissioned its annual industry survey, the same one it had published for years, the same one journalists and analysts quoted as if it were peer-reviewed research. The polling company, a reputable outfit with a long client list, claimed to have surveyed a specific number of C-suite executives in a specific sector in a specific geography. Our contributor knew that sector and geography well enough to count the actual population. The numbers did not work. There were not enough people matching the claimed demographics in the claimed location for the sample to exist.
So our contributor did something nobody was supposed to do. He contacted several of the people who should have been in that sample, people he knew personally, people who fit the exact profile. None of them had been contacted. None had received an outreach. None knew the survey existed.
When confronted, the polling company retreated behind professional discretion. Their methodology was proprietary. Their respondent lists were confidential. They could not disclose who they had spoken to or how they had assembled the panel. They did not deny the accusation directly. They simply declined to address it.
Our contributor raised this with the firm’s marketing leadership. The response was three words: leave it alone.
The survey published on schedule. The press release went out. Journalists cited it. Clients referenced it in board papers. An analyst at a bank used it in a sector report. The numbers entered the information supply chain and became, for all practical purposes, true.
We wish this were unusual. Across 170 combined years inside these firms, not one of us can identify a single survey or report in the last fifteen years that followed every step of the process it claimed to follow. Not one.
How the sausage actually gets made
The public imagines that a Big 4 or MBB report begins with a question and arrives at an answer through research. The actual sequence runs in the opposite direction.
A partner or a marketing team decides what the report should say. The conclusion comes first. It is usually a market-sizing figure that positions the firm’s service line favourably, or a trend narrative that creates urgency for the type of engagement the practice sells. The research is then commissioned to produce evidence that supports the predetermined conclusion.
This is not a secret within the firms. The partners who lead these projects would describe it differently. They would say they are “testing a hypothesis” or “validating a market perspective.” The distinction between testing a hypothesis and confirming a conclusion is supposed to be methodological rigour. In practice, the rigour is absent.
We watched it work the same way across multiple firms, multiple geographies, multiple decades. The pattern was the same everywhere.
Start with the sample. When survey responses came back and the data contradicted the messaging, the team would identify “outliers” and remove them. The justification was always ad hoc and always arrived after the inconvenient data appeared. Nobody decided in advance which responses would constitute outliers. The decision was made retroactively, by the people who needed the data to say something specific.
Then the mathematics got creative. One of the most common abuses we witnessed was averaging percentages. If 30% of respondents in one segment and 50% in another reported a particular behaviour, the report would state that “40% of executives” exhibited that behaviour, regardless of segment size. When one of our contributors flagged this as statistically invalid (you cannot average percentages across groups of different sizes without weighting), the response was genuine confusion. The people assembling the report did not understand why it was wrong. They had been doing it for years.
The questions did the rest. Survey questions were structured to produce the desired result. Leading phrasing, anchoring effects, forced-choice options that excluded inconvenient answers. A question like “How significant is AI to your organisation’s strategy?” with options ranging from “significant” to “extremely significant” will reliably produce a headline claiming that “95% of executives consider AI significant to their strategy.” The option of “not significant at all” was either absent or buried.
And through all of this, the teams went through the motions of analytical process. They ran models. They produced charts. They wrote methodology sections. What they produced was, in most cases, no better than fabrication wearing a lab coat. But the appearance of rigour served a purpose: it gave the marketing team something to point to when anyone asked how the numbers were derived.
The supply chain of manufactured confidence
The polling company incident our contributor described was not an aberration. It sits inside a supply chain that has financial incentives at every link to produce the answer the client wants.
Consider how the commissioning works in practice. Firms hire external polling companies. Those polling companies know that repeat business depends on delivering results the client can use. A polling company that consistently returns data contradicting the firm’s preferred narrative does not get hired again. A polling company that reliably delivers clean, quotable numbers supporting the firm’s thesis gets a multi-year contract.
This incentive structure is not unique to consulting. But consulting adds a complication: the firms commission the research, control the questions, shape the sample criteria, and then publish the results under their own brand with no external review. The polling company provides plausible deniability. The firm can say the research was “conducted by an independent third party” while having determined the outcome before the first question was asked.
In April 2025, the U.S. Department of Justice indicted eight people connected to market research firms Op4G and Slice for a decade-long scheme to fabricate survey data worth $10 million. The defendants recruited “ants” who posed as legitimate survey respondents, were coached on how to answer screener questions, and used VPNs to conceal their locations. Their clients included Google and Seattle Children’s Hospital. Healthcare decisions may have been based on entirely fabricated data.
That case is the criminal extreme. But the academic evidence suggests the everyday reality is not dramatically better. Researchers Noble Kuriakose and Michael Robbins analysed over 1,000 public datasets from international surveys and found that roughly one in five failed a statistical test for fabricated data. NORC at the University of Chicago reported in early 2026 that 40 percent of nonprobability survey interviews in 2025 were likely fraudulent, driven by click farms, bots, and professional survey takers. The Insights Association has called survey fraud an “existential threat” to the industry.
The CDC bleach study illustrates what happens at the other end of this pipeline. In 2020, the CDC reported that Americans were ingesting household cleaners to prevent COVID-19. The finding was reported by over 150 news outlets. When researchers replicated the study with proper quality screening, they found that 100% of the reported bleach ingestion came from problematic respondents: bots, inattentive clickers, and professional survey takers who answered “yes” to questions like “have you ever suffered a fatal heart attack?” The alarming public health finding was entirely an artefact of bad data from an unvetted panel.
Nobody at the CDC fabricated anything. They contracted a market research vendor, the vendor sourced from a common panel supplier, and nobody checked whether the respondents were real. The same supply chain that fills Big 4 survey panels.
The reports get worse. The citations keep coming.
In 2024, accounting professors Jeremiah Green and John Hand published a replication study of McKinsey’s “Diversity Wins” research in Econ Journal Watch. McKinsey’s series of reports, published in 2015, 2018, and 2020, claimed a correlation between executive diversity and financial outperformance. The studies were cited by BlackRock, the Nasdaq, and U.S. government agencies. Green and Hand could not reproduce the claimed correlations. They identified a reverse-causality flaw in McKinsey’s design: McKinsey measured diversity after the performance period, not before. McKinsey refused to share its datasets. When asked, the firm said it stood by its findings.
The diversity studies are worth examining not because diversity is unimportant (it may well drive outperformance through mechanisms McKinsey did not test), but because they illustrate how consulting research enters the policy bloodstream. A non-peer-reviewed study, produced by a firm with a commercial interest in diversity consulting, using a methodology that independent academics could not replicate, influenced board composition rules at major stock exchanges and investment criteria at the world’s largest asset manager. At no point did anyone with the authority to mandate these changes require the underlying data.
McKinsey’s market-sizing projections follow a similar pattern. The firm’s Quantum Technology Monitor projected in April 2026 that quantum computing could generate $400 to $600 billion in value for financial services alone by 2035. A detailed critique published at postquantum.com (which we came across independently and recommend reading in full) showed that the arithmetic on McKinsey’s own slide cannot be reproduced from the assumptions displayed on it, that the impact estimates measure quantum computing and AI together without allocating between them, and that the finance chapter ignores peer-reviewed resource estimates, including two papers from Goldman Sachs’s own researchers, showing the hardware requirements are orders of magnitude beyond what the industry expects to deliver in that timeframe. The author’s conclusion: a headline figure whose published inputs produce a different figure is a graphic, not a forecast, and nobody should commit capital on its basis.
These are “value-at-stake” estimates, a construction McKinsey uses frequently. The phrase sounds technical. What it means is: “this is the total amount of money that could theoretically be affected if every assumption we made turns out to be correct.” It is not a forecast. It is not a projection. It is a ceiling so high that the actual outcome will always fall somewhere beneath it, and the firm can always claim the estimate was directionally correct.
The consulting firms produce these numbers for industries where they sell advisory services. McKinsey publishes AI market-sizing reports while selling AI consulting engagements. The Big 4 publish cybersecurity threat reports while selling cybersecurity services. The conflict of interest is so consistent and so visible that the only explanation for its persistence is that nobody with enough influence cares to challenge it.
Then came the AI reports
If the pre-AI era involved humans cutting corners on research methodology, the AI era has introduced a new efficiency: letting the machine fabricate the citations while the humans fabricate the conclusions.
In 2025 and 2026, GPTZero and academic researchers exposed AI-fabricated content in reports from three of the four Big 4 firms. Deloitte Australia refunded part of a A$440,000 government contract after its welfare-compliance report was found to contain fabricated citations, non-existent academic papers, and a made-up Federal Court quote. Deloitte Canada’s $1.6 million healthcare review for Newfoundland and Labrador contained false academic citations and researchers misattributed to studies they never wrote. EY Canada withdrew a cybersecurity report after GPTZero found that 16 of 27 cited sources were fabricated, including a citation to a McKinsey report that does not exist. KPMG pulled a customer-experience report in which only 5 of 45 citations were accurate.
The irony is almost too precise. These firms sell AI transformation to their clients. They charge millions for advice on how to implement AI responsibly. And they cannot manage the integrity of their own AI-generated content.
But here is the part the coverage misses: the AI reports are not a new category of failure. They are the old failure accelerated. The methodology was already broken. The citations were already decorative. AI simply made the decoration faster and the fabrication more obvious. A report that starts from its conclusion and works backwards does not become honest because a human, rather than a language model, selects the supporting references. It just becomes harder to catch.
Goodhart’s Law with a consulting fee
A reader named Claudine Cassar recently shared an account that captures a parallel phenomenon. KPMG’s U.S. advisory business reportedly told staff to register AI use on 75% of business days. The predictable result: employees started typing random questions into the system, summarising emails they had already read, and automating prompts simply to meet the usage target. The dashboard showed AI “usage” rising. Press releases followed. The metric created the appearance of adoption while revealing nothing about whether the technology improved anyone’s work.
This is the same mechanism applied to different content. The survey that starts from its conclusion. The market-sizing estimate that defines its terms broadly enough to guarantee an impressive number. The AI adoption metric that measures keystrokes instead of outcomes. EY’s definition of “AI-related revenue” is instructive here: the firm counts everything from enterprise-wide transformations to AI governance frameworks, a category broad enough that adding a chatbot module to a broader programme makes the entire engagement “AI-led.” When you have spent $2.4 billion and need to show a return, every practice leader has an incentive to relabel existing work as AI-adjacent.
The reported numbers are often technically defensible and practically meaningless. “90% of EY professionals have completed foundational AI training” sounds impressive until you learn the metric is “awarded or initiated,” not completed. “161,000 AI badges” could mean 161,000 people who clicked through an online module. We sat in the partner meetings where these metrics get constructed. The pressure to demonstrate ROI on AI investment is real. The easiest way to demonstrate ROI is to redefine what counts as AI revenue until the growth rate looks impressive enough for the press release.
What to do when you read the next one
We are not suggesting that every number in every Big 4 or MBB report is fabricated. Some of the underlying research is competent. Some of the analysis is worth reading. The problem is that there is no reliable way for an outside reader to distinguish the competent research from the manufactured kind, because the same brand, the same formatting, and the same authoritative tone are applied to both.
These reports are, in library-science terms, “grey literature,” material produced outside traditional academic or commercial publishing that is not subject to peer review. They acquire authority through repetition: a McKinsey estimate gets cited by a journalist, picked up by an analyst, referenced in an academic paper, and trained into a language model. By the time it reaches a board presentation, the number has been laundered through enough credible-seeming intermediaries that nobody traces it back to a marketing team that decided what the report should say before the research began.
So treat them accordingly.
When a Big 4 or MBB report sizes a market, check whether the firm sells services into that market. If it does, the estimate is advertising. Read the methodology section, if one exists, and check whether the stated assumptions actually produce the headline number. (In the McKinsey quantum case, they do not.) When a report claims survey results, look for the sample size, the sampling method, the response rate, and the exact question wording. If any of these are missing, the results are decoration.
And if you see a statistic from a consulting firm cited in an academic paper, a news article, or a board pack, ask one question: has anyone outside the firm that produced it attempted to reproduce the finding? If the answer is no, the number has never been tested. It has only been repeated.
We built these reports. We sat in the rooms where the conclusions were written before the research was commissioned. We watched polling companies deliver impossible samples and marketing teams suppress inconvenient data. We averaged the percentages. We wrote the methodology sections that nobody read. We knew, and we said nothing, because the reports served their purpose and the purpose was never truth.
The purpose was always the same: give the market a number it could cite, attach our brand to it, and wait for the phone to ring.
What is the worst example of a manufactured number you have seen in a Big 4 or MBB report? If you work at one of these firms now, what does the process actually look like from the inside? The comments are open, and so is the inbox.


