mrkeyoor.com_
Wed 02 Sept 08:46 UTC
AI6 min read

Dan Luu's 638-Point Audit Tests Ed Zitron's AI Forecasts

A viral prediction scorecard catches clear misses in Gemini users and Big Tech growth, while exposing how hard it is to grade broad claims about AI capability.

A live snapshot at 07:30 UTC on September 2 put a new critique of AI skepticism at 638 Hacker News points and 679 comments. That unusually busy discussion matters because its subject is measurement: when a writer makes repeated, confident calls about an AI crash, what evidence should count when readers check the record?

Engineer and essayist Dan Luu supplied the match. In his audit of Ed Zitron's predictions, Luu follows claims made between February 2024 and November 2025 about model progress, OpenAI's growth, Gemini adoption and the health of large technology companies. He judges nearly every resolved claim wrong. The dated trail of claims is the useful result; readers can inspect it without joining either AI camp.

For developers, this is more practical than the usual argument over whether AI is overhyped. Infrastructure plans, hiring and product bets depend on time-bound expectations. Zitron's February 2024 essay said generative AI was reaching the upper limits of what it could do; by February 2025, he was calling Google's 500 million-user Gemini target unrealistic. Those are different kinds of forecast, as the original February 2024 post makes plain, and they need different tests.

The audit has a missing denominator

Luu discloses an important part of his method near the end of the essay. He asked ChatGPT to produce a list of predictions, read or skimmed the linked posts, removed cases where the model had misread the text, and excluded statements he considered unfalsifiable or tautological. He also says ChatGPT and Claude later found minor errors in his draft, which he corrected before publication. That account is unusually candid, but the published audit does not include a complete corpus, a selection protocol or a numerical score.

The omission does not erase the misses Luu documents. It limits the larger claim that the list captures Zitron's prediction record as a whole. A reproducible scorecard would define the search period, archive every candidate statement, record exclusion reasons and let a second reviewer grade each result without seeing the first grade. Luu acknowledges the workload: his method note says a stronger version would likely need several independent raters.

This distinction surfaced in the Hacker News thread. Commenters argued over whether Luu had sampled only failures, whether being early should count as wrong, and whether AI promoters deserve the same treatment. Those comments are reactions rather than evidence, but their volume shows why the 638-point discussion took off: people lack a shared rule for scoring technology forecasts after the deadline passes.

Public accounts beat broad diagnoses

Luu's strongest material uses company accounts against a sweeping corporate claim. He points to a November 2024 talk in which Zitron described Meta as a dying company and grouped Google and Microsoft into an ecosystem that no longer knew how to grow. Luu then compares that diagnosis with reported revenue and operating income, laying out the figures in calendar-year tables.

Meta's latest quarter makes the narrow verdict straightforward. Revenue reached $60.8 billion in the second quarter of 2026, up 28 percent from a year earlier, while first-half revenue rose from $89.8 billion to $117.1 billion. Yet the same Meta earnings release contains the pressure that a fair scorecard should retain: quarterly operating income fell 8 percent, net income fell 14 percent, and capital expenditure reached $31.1 billion. Meta was growing, though that growth came with a much heavier cost base.

Alphabet's full-year numbers are harder to square with a near-term death diagnosis. It reported $402.8 billion in 2025 revenue, up 15 percent, and $129.0 billion in operating income, up from $112.4 billion. Google Cloud revenue rose 48 percent in the fourth quarter. Alphabet also said the Gemini app had passed 750 million monthly active users, according to its February 2026 earnings release.

Microsoft ended its 2026 fiscal year with $331.8 billion in revenue and $155.2 billion in operating income, increases of 18 percent and 21 percent. Intelligent Cloud revenue grew 32 percent in the June quarter, while the More Personal Computing segment declined 4 percent. The company's full-year release therefore supports Luu on overall growth while preserving a detail his summary table flattens: the businesses inside Microsoft were moving at very different speeds.

These consolidated results rebut the claim that the companies were already dying or unable to grow. They do not settle the return on AI spending. Meta still earns almost all of its revenue from advertising, Alphabet does not report Gemini as a standalone segment, and Microsoft's accounts combine AI services with much larger cloud and productivity businesses. The company disclosures can disprove a collapse call without proving that every GPU purchase will earn an adequate return.

Gemini gives the scorecard a clean test

In February 2025, Zitron wrote that Sundar Pichai's target of 500 million Gemini users by the end of that year was so unrealistic that someone at Google should be fired. The line appears in There Is No AI Revolution, alongside an argument that Google's distribution had failed to produce meaningful consumer adoption. It has a named metric and a deadline, which makes it much easier to grade than a prediction that a technology has peaked.

Alphabet later reported more than 750 million monthly active users for the Gemini app. On the stated metric, the forecast missed by at least 250 million users. The result says less about depth of use, retention or profit because Alphabet's earnings document gives the monthly active figure without a standalone Gemini income statement. A distribution push could help explain the count, but it cannot make 750 million smaller than the rejected 500 million threshold.

That example also reveals why good forecasts need units. A critic could reasonably ask how many Gemini users arrive voluntarily, how often they return, what inference costs Google bears, or whether the app produces incremental revenue. Zitron's published claim chose monthly users instead. Once the outcome arrived, switching the test to profit or user intent would evaluate a new forecast.

Capability forecasts still need definitions

The capability claims are less tidy. Zitron wrote in February 2024 that generative AI was nearing its upper limits and argued in April that the industry was running out of usable training data. The April essay connected data scarcity and synthetic-data degradation to a claim that models could not progress far beyond their then-current state. Luu marks these calls wrong because later systems improved.

There is clear tension between a hard ceiling and subsequent model gains, but neither writer establishes one stable yardstick for the whole period. Accuracy can mean fewer factual errors, better coding, stronger visual consistency or performance on a chosen evaluation. Luu even notes that many widely cited AI evaluations are flawed and says one prominent progress benchmark is not meaningful. His audit is persuasive when it identifies a dated contradiction, and less conclusive when the disputed noun is simply "progress."

A stronger capability forecast would name a task distribution, model-access rules, cost ceiling and resolution date. It might predict, for example, that no publicly available model under a stated price will exceed a specified score on a frozen test by December 2027. Luu says he records confidence levels for his own predictions and points readers toward proper scoring rules; that practice, described in the method appendix, makes repeated certainty expensive when outcomes disagree.

Capital spending sets the next test

The financial argument now has dates attached to it. Alphabet said it expected $175 billion to $185 billion in 2026 capital expenditure, while Meta narrowed its forecast to $130 billion to $145 billion. Those commitments appear in the companies' Alphabet and Meta releases, and they keep the underlying bubble question open even after several short-term collapse predictions failed.

Watch the next filings for AI-linked revenue disclosure, depreciation, free cash flow and any retreat from those spending plans. Also watch whether future prediction audits publish their full sample before assigning a win-loss record. The 638-point response rewarded a pile of receipts; the unresolved question is whether critics and boosters will start writing forecasts that can be scored without first arguing over what they meant. Luu's essay shows both the value of checking and the cost of leaving the rules until after the result.

Sources

  1. How accurate have Ed Zitron's AI skeptic predictions been?
  2. Hacker News discussion: How accurate have Ed Zitron's AI skeptic predictions been?
  3. Subprime Intelligence
  4. Bubble Trouble
  5. There Is No AI Revolution
  6. Alphabet Announces Fourth Quarter and Fiscal Year 2025 Results
  7. Meta Reports Second Quarter 2026 Results
  8. Microsoft Fiscal Year 2026 Fourth Quarter Results