Nothing.
This is the environment file as it stands today, not a record of this run. The run predates environment capture, so what it actually had on hand was never recorded and cannot be recovered.
# Run 0 — find the areas worth mapping Paste this whole thing. Nothing to fill in. --- ## What this run is for I want to spend future nights mapping open questions in a technical field. I don't know which field yet, and I don't want to pick from memory or from a list someone generated without checking. Your job is to go find the candidates and show me the evidence. ## What makes an area a good candidate Look for areas where **claims are loud and independent verification is thin**. Concretely, any of these: - **A live dispute.** Two credible parties publicly disagree about whether a result holds, and it hasn't been settled. Failed replications, published rebuttals, comment-and-response exchanges, editorial notes. - **A loud claim nobody independently checked.** A widely cited result whose data or code isn't available, or where every follow-up comes from the original group. - **Nobody's job.** The question falls between fields, or everyone qualified to settle it has a commercial or reputational stake. An area is a *bad* candidate if it's simply hard and everyone agrees it's hard. Difficulty isn't the signal. Contested-ness is. ## What to produce Around ten candidate areas. Fewer is fine if that's what's real. For each one: - **The area**, in a sentence a non-specialist understands. - **What's contested** — the specific claim in dispute, not the general topic. "Whether method X actually beats baseline Y" beats "AI in chip design." - **Who's on each side**, and what their stake is if they have one. - **The evidence this is live**: at least two sources, with titles, authors, years, and links. At least one should be a document where someone states the disagreement directly — a rebuttal, a failed replication, a comment paper, an editorial note. Quote or closely paraphrase the sentence that shows it, and say where in the document it appears. - **How long it's been unresolved.** - **Whether a careful outsider could actually map it** — is the evidence public, or is the key data proprietary and the whole thing unmappable from outside? Say which. - **Your confidence** that this is genuinely unresolved rather than settled somewhere you didn't look: high, medium, low. ## Where to look Cast wide. Include at least one candidate from each of: semiconductors and hardware, AI/ML internals, quantum, and somewhere none of those — biology, climate, materials, economics, medicine. Do not restrict yourself to fields I might be expected to care about; a good candidate in an unexpected field is more valuable than a mediocre one in a familiar field. Useful places for this kind of thing: PubPeer, Retraction Watch, arXiv comment threads and v2+ revisions, journal comment-and-reply sections, replication studies, reproducibility workshops, editorial expressions of concern. ## Rules - Every claim about a dispute needs a source. If you can't find someone publicly saying it, it doesn't go in the list. - Don't include a controversy that's already been resolved — check whether the argument continued after the last exchange you found. - Don't pad. Six real candidates beat ten with four invented. - Rank them at the end: which would make the best first mapping run, and why. Say plainly if you think some are not worth doing. ## Output Write `output.md` in this run's folder. Then a section called **Notes on this run**: what you searched, what you couldn't reach, what you're unsure about, and anything that surprised you. Blunt. This part is for me, not for a reader.
**The area.** Whether Google's deep-reinforcement-learning method for placing the big blocks (macros) on a chip — published in *Nature* in 2021, later branded AlphaChip — actually produces better layouts than a well-tuned classical algorithm and human designers, or whether the headline result falls apart under a fair, reproducible comparison. **What's contested (the specific claim).** Not "is ML useful in chip design." The sharp claim in dispute: that AlphaChip's RL beats (a) a properly tuned simulated-annealing baseline and (b) commercial tools / human designers on comparable compute. Google says yes and that critics ran the method wrong (no pre-training, ~20× less compute, not trained to convergence). Critics say the original *Nature* comparison wasn't reproducible and omitted steps needed to check it. - *Defenders:* Anna Goldie, Azalia Mirhoseini, Jeff Dean (Google/DeepMind). Authors of the original paper; AlphaChip is a flagship Google "AI-for-chip-design" result taped out in TPUs — reputational and commercial stake. - *Critics:* Andrew Kahng and Chung-Kuan Cheng's group (UC San Diego / TILOS), authors of the "Assessment of Reinforcement Learning for Macro Placement" line; and Igor Markov (long-time EDA figure), author of "The False Dawn," who also aired integrity allegations. - **[double-checked]** Igor L. Markov, *"The False Dawn: Reevaluating Google's Reinforcement Learning for Chip Macro Placement,"* arXiv:2306.09633 (2023; published in *Communications of the ACM*, 2024). https://arxiv.org/abs/2306.09633 — abstract, as surfaced: *"Crosschecked data indicate that the integrity of the Nature paper is substantially undermined owing to errors in conduct, analysis and reporting,"* and *"Before publishing, Google rebuffed internal allegations of fraud, which still stand."* Directly states the dispute. - A. Goldie, A. Mirhoseini, J. Dean, *"That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design,"* arXiv:2411.10053 (Nov 2024). https://arxiv.org/abs/2411.10053 — argues the skepticism is "unfounded" and that Cheng et al. "did no pre-training … used [~20×] less compute, and did not train to convergence." The rebuttal. - Nature, *"Addendum: A graph placement methodology for fast chip design,"* Nature 634, E10–E11 (2024), DOI 10.1038/s41586-024-08032-5. https://www.nature.com/articles/s41586-024-08032-5 — Nature attached an editor's note (Sept 2023) that "the paper's performance claims had been called into question," then published an Addendum rather than a correction. - Cheng et al., *"An Updated Assessment of Reinforcement Learning for Macro Placement,"* arXiv:2302.11014, **accepted in IEEE TCAD in Dec 2025** — i.e., the critics got a peer-reviewed venue *after* Google's rebuttal and the Addendum. https://arxiv.org/abs/2302.11014 . Original paper: *Nature* 594, 207 (2021), https://www.nature.com/articles/s41586-021-03544-w . **How long unresolved.** ~4 years (paper mid-2021; critiques 2022–2023; Addendum + Google rebuttal late 2024; critics' peer-reviewed update Dec 2025). Still moving. **Mappable from outside?** The most mappable of the hardware candidates. Original code (`circuit_training`) and the TILOS/MacroPlacement benchmark repo are public; both sides' papers are public. What an outsider *cannot* touch: Google's internal pre-training corpus, its production placements, and the exact compute used — so you can adjudicate the reproducible-benchmark question but not the "it works in shipping TPUs" claim. **Confidence genuinely unresolved: HIGH.** Both sides published in 2024–2025; the Addendum did not make critics stand down; no consensus exists. ---
**The area.** Whether Microsoft's InAs–Al hybrid nanowire devices actually host Majorana zero modes and constitute a "topological qubit" (the Feb-2025 "Majorana 1" announcement), or whether the signatures are trivial disorder artifacts.
**What's contested (the specific claim).** Two linked claims: (a) that Microsoft's "topological gap protocol" (TGP) reliably identifies a topological phase without false positives, and (b) that the 2025 single-shot parity-measurement device demonstrates a topological qubit. Critic Henry Legg argues the TGP is ill-defined and that the readout happened in regions that look gapless/disordered; Microsoft calls the objections minor bugs / a straw man.
- *Pro:* Microsoft Quantum; spokesperson Chetan Nayak (Technical Fellow). Enormous stake — "Majorana 1" is Microsoft's entire quantum roadmap and a public product announcement; DARPA benchmarking involvement.
- *Con:* Henry F. Legg (University of St Andrews), plus condensed-matter physicists who voiced doubt at the March-2025 APS meeting. Legg is an independent academic (relatively low commercial stake — notable, given this field's "everyone qualified has a stake" pattern).
- **[double-checked]** Microsoft Quantum, *"Interferometric single-shot parity measurement in InAs–Al hybrid devices,"* Nature 638, 651–655 (2025), DOI 10.1038/s41586-024-08445-2. https://www.nature.com/articles/s41586-024-08445-2 — Nature's editors appended an editorial note (as surfaced): *"the results in this manuscript do not represent evidence for the presence of Majorana zero modes in the reported devices."* An editorial disclaimer on the published paper is a strong "live" signal.
- **[double-checked]** Henry F. Legg, *"On the robustness of topological gap detection via transport,"* a *Matters Arising* in Nature 654, E22–E26 (2026), DOI 10.1038/s41586-026-10567-8, published ~24 June 2026, **with a Microsoft Reply in the same issue.** https://www.nature.com/articles/s41586-026-10567-8 — Legg's point (as surfaced): the claimed parity readout "occurred in regions of phase space with considerable disorder that appear gapless," and shifting measurement windows flips the TGP's own verdict on the same device region. Coverage: The Register, *"Boffin claims Microsoft's supposed quantum leap does not compute due to 'basic Python errors'"* (24 Jun 2026), https://www.theregister.com/research/2026/06/24/boffin-claims-microsofts-supposed-quantum-leap-does-not-compute-due-to-basic-python-errors/5260489 ; TechXplore, *"Critique challenges Microsoft's quantum computing claims"* (Jun 2026), https://techxplore.com/news/2026-06-microsoft-quantum.html .
- Foundational preprint comments: Legg, arXiv:2502.19560 (Feb 2025), https://arxiv.org/abs/2502.19560 ; and arXiv:2503.08944 (Mar 2025), https://arxiv.org/abs/2503.08944 .
**How long unresolved.** The Majorana-signature dispute in these devices traces to the retracted 2018 Zhang et al. *Nature* paper; the current TGP/qubit round runs 2023–2026, with the most recent formal *Nature* exchange in **June 2026** — two months before this run. Actively live.
**Mappable from outside?** Partly. The TGP logic and released transport data are analyzable from public documents — Legg re-analyzed released data and pointed at specific code behaviors. But the definitive raw device data and fabrication details are Microsoft-held, and the physics question ("is there a Majorana?") needs experiments only Microsoft runs at scale. You can map the *dispute* thoroughly; you cannot settle the *physics*.
**Confidence genuinely unresolved: HIGH.** A *Nature* Matters-Arising + Reply in June 2026 with neither side conceding.
---
**The area.** D-Wave claims its Advantage2 annealer performed a "beyond-classical" simulation of spin-glass quench dynamics that no classical computer can match in reasonable time. **What's contested (the specific claim).** Whether classical tensor-network methods reproduce D-Wave's results on ordinary hardware. Challengers say yes — they matched the physics "on a laptop." D-Wave says no — the classical methods fail on the *hardest* instances and highest-order measurements, so the supremacy claim stands. - *D-Wave Quantum Inc. (NYSE: QBTS):* a **publicly traded company** whose product narrative and stock lean on the claim; it filed an SEC 8-K defending it. Original paper: A. King et al. - *Challengers:* Joseph Tindall, A. F. Mello, M. Fishman, E. M. Stoudenmire, D. Sels (Flatiron Institute / Simons Foundation + Boston University). Elite computational-physics reputations at stake; *Science* published both sides. - D-Wave original: King et al., *"Beyond-classical computation in quantum simulation,"* Science (2025), DOI 10.1126/science.ado6285, https://www.science.org/doi/10.1126/science.ado6285 ; preprint arXiv:2403.00910. - **[double-checked]** Tindall et al., *"Dynamics of disordered quantum systems with two- and three-dimensional tensor networks,"* arXiv:2503.05693, published in *Science* (2026, DOI 10.1126/science.adx2728). https://arxiv.org/abs/2503.05693 — abstract, as surfaced: these experiments "were claimed to be beyond the reach of classical computation … we find that state-of-the-art accuracies can be achieved with modest computational resources." Publicized by the Simons Foundation, *"Quantum Dynamics Breakthrough Overturns Claim of 'Quantum Supremacy,'"* 21 May 2026, https://www.simonsfoundation.org/2026/05/21/quantum-dynamics-breakthrough-overturns-claim-of-quantum-supremacy-opens-new-research-directions/ . - **[double-checked]** D-Wave's rebuttal, *"D-Wave's Quantum Supremacy Result Stands"* (26 May 2026), https://www.dwavequantum.com/company/newsroom/press-release/d-wave-s-quantum-supremacy-result-stands/ , and SEC 8-K, https://www.stocktitan.net/sec-filings/QBTS/8-k-d-wave-quantum-inc-reports-material-event-04687ff8a7dd.html — position (as surfaced): the "overturned" claim is "inaccurate and not supported by the scientific record" because the new work "does not reproduce the full scope" nor "the hardest problem instances." Technical comment: arXiv:2504.06283. **How long unresolved.** Claim March 2025; classical challenges March 2025 → *Science* publication May 2026; D-Wave counter-rebuttal May 2026. ~15 months and escalating. **Mappable from outside?** Highly. Everything is public — two *Science* papers, an arXiv comment-and-reply, company press releases, even an SEC filing. One of the cleanest live disputes here. ---
**The area.** USTC's photonic Jiuzhang machines claim quantum computational advantage via Gaussian boson sampling (GBS) — sampling from distributions claimed classically intractable. **What's contested (the specific claim).** Whether real, *lossy* GBS experiments actually beat classical computers, or whether photon loss and noise let a classical algorithm match — or even beat — the experiment's own sample quality. Narrowly: a classical sampler that produces samples *closer to the ideal distribution than the noisy experiment itself* would negate the advantage; USTC's newer experiments claim to beat every proposed classical method. - *USTC* (Jian-Wei Pan, Chao-Yang Lu et al.): flagship Chinese quantum-advantage program; national-prestige stake. - *Classical challengers:* Changhun Oh, Minzhao Liu, Yuri Alexeev, Bill Fefferman, Liang Jiang (Chicago / Argonne / KAIST). Stake: credibility of the "advantage" milestone. A running cat-and-mouse. - Oh, Liu, Alexeev, Fefferman, Jiang, *"Classical algorithm for simulating experimental Gaussian boson sampling,"* Nature Physics 20, 1461–1468 (2024), DOI 10.1038/s41567-024-02535-8, https://www.nature.com/articles/s41567-024-02535-8 ; preprint arXiv:2306.03709 — abstract, as surfaced: *"Our classical sampler can simulate the ideal distribution better than the experiment can, which calls into question the claims of experimental quantum advantage."* States the disagreement. - USTC escalation: *"Robust quantum computational advantage with programmable 3050-photon Gaussian boson sampling"* (Jiuzhang 4.0), arXiv:2508.09092 (Aug 2025), https://arxiv.org/abs/2508.09092 — abstract, as surfaced: results "outperform all classical spoofing algorithms, particularly the matrix product state (MPS) method … recently proposed to utilise photon loss to reduce the classical simulation complexity of GBS." Answers the challenge. - Original claim lineage: *"Quantum computational advantage using photons,"* Science (2020), https://www.science.org/doi/10.1126/science.abe8770 . **How long unresolved.** Running since 2020; loss/noise-exploiting classical attacks 2022–2024; USTC's answer 2025. ~5–6 years, still cycling. **Mappable from outside?** Highly — the whole back-and-forth is on arXiv and in Nature/Science/Nature Physics. Weak spot in the *science* (not the mapping): raw experimental sample sets aren't always released, so challengers work from published statistics — part of why it never fully closes. **Confidence genuinely unresolved: MEDIUM-HIGH.** Genuinely ongoing, but the pattern is escalation (each new machine resets the target) rather than a clean two-party fight over one paper. ---
**The area.** Apple's high-profile claim that frontier "reasoning" models suffer a "complete accuracy collapse" past a complexity threshold on planning puzzles is disputed as an experimental-design artifact — and the rebuttals are themselves disputed. **What's contested (the specific claim).** Shojaee et al.'s (Apple) finding that models like Claude, o3, and Gemini 2.5 fail *fundamentally* on Tower of Hanoi / River Crossing beyond a complexity threshold. The rebuttal claims the "collapse" is (a) models hitting output-token limits (and saying so), (b) an autograder scoring truncation as reasoning failure, and (c) River-Crossing instances that are mathematically unsolvable for N>5 — so models were penalized for refusing impossible problems. The counter-claim (Marcus) is that the rebuttals "fall short" and the limitation stands. - *Reasoning-skeptic:* P. Shojaee, I. Mirzadeh et al. (Apple). Apple is a frontier-model laggard; a "reasoning is illusory" result is reputationally convenient — critics note this. - *Rebuttal:* C. Opus & A. Lawsen (Open Philanthropy-affiliated), arguing methodology flaws. - *Defends Apple / attacks the rebuttals:* Gary Marcus. Long-standing public stake in LLM skepticism. - Apple: Shojaee et al., *"The Illusion of Thinking …,"* arXiv:2506.06941 (June 2025), https://arxiv.org/abs/2506.06941 (also https://machinelearning.apple.com/research/illusion-of-thinking ). - **[double-checked]** C. Opus & A. Lawsen, *"Comment on 'The Illusion of Thinking' …,"* arXiv:2506.09250 (2025), https://arxiv.org/abs/2506.09250 — abstract, as surfaced: findings "primarily reflect experimental design limitations rather than fundamental reasoning failures … Tower of Hanoi experiments systematically exceed model output token limits at reported failure points, with models explicitly acknowledging these constraints." The direct rebuttal. - Gary Marcus, *"Seven replies to the viral Apple reasoning paper — and why they fall short,"* June 2025, https://garymarcus.substack.com/p/seven-replies-to-the-viral-apple — the argument continuing *after* the rebuttal. Also *"Rethinking the Illusion of Thinking,"* arXiv:2507.01231 (July 2025). Upstream skeptic line: Apple's *GSM-Symbolic*, arXiv:2410.05229 (ICLR 2025). **How long unresolved.** GSM-Symbolic Oct 2024; *Illusion of Thinking* June 2025; rebuttal + counter-defense within days; counter-rebuttal July 2025. Active, no concession. **Mappable from outside?** Yes, largely — and you could partly *re-adjudicate* it: the papers, rebuttal, and counter-defense are public, the puzzles are reproducible, and the token-limit argument is checkable by anyone with API access. Only semi-hidden piece: the exact model versions / decoding settings Apple used. ---
**The area.** Whether the sudden capability "jumps" with scale that Wei et al. reported are genuine phase transitions, or an illusion created by discontinuous evaluation metrics. **What's contested (the specific claim).** Schaeffer et al. claim emergence "appears due to the researcher's choice of metric" — swap exact-match/accuracy for a continuous metric and the sharp jump smooths out, so emergence isn't a fundamental property of scaling. Wei's camp concedes some tasks smooth under continuous metrics but argues exact-match *is* what we care about for many tasks, and some capabilities remain abruptly emergent even under continuous metrics. Live sub-dispute (2025): whether *any* residual emergence survives once you control for metric and random-seed variance. - *"Mirage":* Schaeffer, Miranda, Koyejo (Stanford). Paper won a NeurIPS 2023 outstanding-paper award — reputational investment in the deflationary claim. - *"Emergence is real":* Jason Wei (then Google) et al. — authors of the claim under attack. - *Keeping it open:* *Random Scaling of Emergent Capabilities*, arXiv:2502.17356 (Feb 2025), argues continuous *loss* can still be bimodal across seeds — i.e., some emergence persists. - Schaeffer, Miranda, Koyejo, *"Are Emergent Abilities of Large Language Models a Mirage?"* NeurIPS 2023, arXiv:2304.15004, https://arxiv.org/abs/2304.15004 — abstract, as surfaced: emergent abilities "appear due to the researcher's choice of metric rather than due to fundamental changes in model behavior with scale." States the disagreement, against Wei et al. - Jason Wei, *"Common arguments regarding emergent abilities,"* 2023, https://www.jasonwei.net/blog/common-arguments-regarding-emergent-abilities — concedes some curves smooth under log-prob/edit-distance metrics but argues hard metrics like exact match are what you actually want. Original target: Wei et al., *"Emergent Abilities of Large Language Models,"* TMLR 2022, arXiv:2206.07682. - Continuation: *Random Scaling of Emergent Capabilities*, arXiv:2502.17356, https://arxiv.org/abs/2502.17356 . **How long unresolved.** Claim 2022; mirage 2023; still generating adjudicating papers in 2025. ~3 years. **Mappable from outside?** Fully. A public argument over metrics and public benchmark curves (BIG-Bench, GPT-3 arithmetic); an outsider can re-plot the same curves under different metrics. No proprietary data needed. **Confidence genuinely unresolved: MEDIUM-HIGH.** Both sides credible and still publishing, but the field is drifting toward a "nuanced middle" — softening, not exploding. ---
**The area.** Whether a foundational 2006 *Nature* result claiming a specific amyloid-beta oligomer (Aβ*56) directly causes memory impairment — a paper that helped steer two decades of Alzheimer's funding toward amyloid — was built on manipulated images. **What's contested (the specific claim).** Narrowly: whether the western-blot images in Lesné et al. 2006 are genuine. The journal and nearly all co-authors say they were spliced/duplicated/erased and can't be verified; the first author says the retraction is wrong. It stays live because the accused author still publicly rejects the finding and objected to a *second* retraction in Jan 2025. - *Retraction side:* Nature editors + co-authors (Karen Ashe, the senior author; Koh, Kotilinek, Kayed, Glabe, Gallagher). Stake: credibility of the broader amyloid field. - *Dissent:* Sylvain Lesné (first author), who disputes the retraction and denies doctoring images. Stake: his career (resigned his tenured UMN post, effective March 2025) and ~20 of his papers under scrutiny. - *Nobody's-job angle:* the detection came from independent sleuths (Matthew Schrag, Elisabeth Bik), not the funders or the journal. - **[double-checked]** *"Retraction Note: A specific amyloid-β protein assembly in the brain impairs memory,"* Nature (2024), DOI 10.1038/s41586-024-07691-8, https://www.nature.com/articles/s41586-024-07691-8 — the co-authors listed "agreed with the retraction … Sylvain Lesné disagreed with the retraction." The retraction shipped over the first author's explicit objection — that single line is the live dispute. (Confirmed near-verbatim on a second independent search.) - *"Alzheimer's scientist resigns after university finds 'data integrity concerns' in papers,"* Science/AAAS (2025), https://www.science.org/content/article/alzheimer-s-scientist-resigns-after-university-finds-data-integrity-concerns-papers — Lesné resigned from UMN effective March 2025 after a university investigation; he also objected to a separate Science Signaling retraction (Jan 2025). The argument continued past the 2024 retraction. - Background: Alzforum, *"Sylvain Lesné, Who Found Aβ*56, Accused of Image Manipulation,"* https://www.alzforum.org/news/community-news/sylvain-lesne-who-found-av56-accused-image-manipulation ; original (retracted) paper: Nature 440, 352 (2006), https://www.nature.com/articles/nature04533 . **How long unresolved.** Allegations public July 2022; 2006 paper retracted June 2024; author still disputing and further retractions landing through 2025. ~3–4 years, open. **Mappable from outside?** Yes — richly. PubPeer threads, the retraction notices, *Science*'s reporting, and the ongoing investigations are all public. One of the best-documented cases available. **Confidence genuinely unresolved: HIGH.** Caveat: some read this as a fraud story that is *resolving against* Lesné rather than a symmetric scientific dispute. It is genuinely unresolved as an *argument* (the author has not conceded), but if you want a two-sides-both-plausible fight, this is more one-sided than, say, D-Wave. ---
**The area.** Whether the tumor-shrinkage and progression-delay measures that now underpin most FDA cancer-drug approvals reliably translate into patients living longer or better. **What's contested (the specific claim).** That treatment effects on surrogate endpoints (progression-free survival, objective response rate) are *validated* predictors of overall-survival benefit strong enough to justify approval/prescribing. Trial-level meta-analyses reach opposing conclusions — some report strong PFS↔OS correlation in specific first-line settings, others find correlations "generally low" across oncology — and the camps have not converged. - *Skeptics:* Vinay Prasad & Chul Kim (2015); Bishal Gyawali. Stake: patient protection; identities built on the critique. - *Defenders/users:* FDA, NCCN guideline-writers, industry sponsors. Stake: approval speed, revenue, regulatory throughput. - *Live twist:* Prasad — the field's most cited *critic* of surrogate-based approval — became FDA CBER Director and is now reported defending surrogates from inside the regulator. The disagreement is now partly *within one person's own record*, while external critics (Gyawali) continue. - **[double-checked]** Prasad V, Kim C, Burotto M, Vandross A, *"The Strength of Association Between Surrogate End Points and Survival in Oncology: A Systematic Review of Trial-Level Meta-analyses,"* JAMA Internal Medicine 175(8):1389–1398 (2015), https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2323416 (PMID 26098871) — finds the correlations "generally low" and the evidence "limited." The loud skeptic claim. - *"Surrogate end points in oncology: the speed–uncertainty trade-off …,"* Nature Reviews Clinical Oncology (2025), DOI 10.1038/s41571-025-01007-z, https://www.nature.com/articles/s41571-025-01007-z — frames surrogate use as an unresolved trade-off, not settled science (2025). - Continued tension: *"US FDA's Prasad: 'We Will Always Embrace Surrogate Endpoints,'"* Citeline/Pink Sheet (2025) — the pro-surrogate stance in the regulator's own words. **Title-confirmed only** (body paywalled + egress-blocked). **How long unresolved.** Actively disputed since at least 2015; conflicting meta-analyses and the Prasad-at-FDA reversal keep it hot in 2025. ~10+ years. **Mappable from outside?** Yes — one of the most publicly documented methodological disputes in medicine (dueling meta-analyses, JAMA/NRCO commentary, FDA statements). The individual patient-level trial datasets that would *settle* the correlations are often proprietary to sponsors — a clean "loud claim, verification thin" feature. **Confidence genuinely unresolved: HIGH** it's contested and live. **Medium** only on the Prasad "always embrace" quote (title-level). ---
**The area.** Whether a drop in an epigenetic "aging clock" reading after an intervention (partial reprogramming, plasmapheresis, senolytics) is real biological-age reversal or a measurement artifact that doesn't track actual function. **What's contested (the specific claim).** The claim — loud in longevity biotech — that observing a lower clock age after an intervention *demonstrates* rejuvenation. Critics argue the clocks can't validate this (the "true biological age" of a reprogrammed cell is unknown, so clock predictions are unfalsifiable), that clock reversal can occur without any age-related functional change, and that a clock may register "improvement" merely by suppressing beneficial repair. - *Skeptics:* Kriukov et al. 2024 (associated with Vadim Gladyshev's group, Harvard); a UC San Diego group arguing clock changes are a symptom, not the cause, of aging. - *Proponents:* reprogramming/longevity labs and companies, plus consumer "biological age" test firms. Large commercial stake — valuations, trial endpoints, and products all lean on clocks being valid readouts. - *Nobody's-job angle:* nearly everyone qualified to adjudicate built the clock, sells the test, or runs the reprogramming company. - Kriukov D. et al., *"Epistemic uncertainty challenges aging clock reliability in predicting rejuvenation effects,"* Aging Cell (2024), DOI 10.1111/acel.14283, https://onlinelibrary.wiley.com/doi/10.1111/acel.14283 — as surfaced: aging clocks are used to validate rejuvenation during reprogramming, but "their predictions are unverifiable due to the unknown true biological ages of reprogrammed cells." The disagreement, stated against the field's dominant practice. - UC San Diego, *"Why Our Biological Clock Ticks"* (2024), https://today.ucsd.edu/story/why-our-biological-clock-ticks-research-reconciles-major-theories-of-aging — as surfaced: institutions and companies "are betting on turning back the epigenetic clock … but our research suggests that this may only be treating a symptom of aging, not the underlying cause." - Supporting: Aging-US, *"Why Epigenetic Clocks May Fail to Measure Anti-Aging Effects,"* https://www.aging-us.com/news-room/why-epigenetic-clocks-may-fail-to-measure-anti-aging-effects . **How long unresolved.** Escalated 2022–2024 as partial-reprogramming and clock-based claims proliferated; the peer-reviewed pushback (Kriukov 2024) is recent and unanswered at the field level. Open in 2025. **Mappable from outside?** Mostly — published critiques + the promotional claims they target are public. Caveat: some of the loudest proponent claims (e.g., non-human-primate reprogramming "signals") aren't peer-reviewed and sit inside companies, so part of the proponent side is proprietary. **Confidence genuinely unresolved: MEDIUM.** Real, published, unrebutted critique from two independent groups — but a diffuse methodological dispute (many parties, no single comment-and-reply) rather than one crisp fight. ---
**The area.** About a fifth of the newest generation of climate models have very high sensitivity (equilibrium warming >5 °C for doubled CO₂) and "run hot" vs. observed warming; the field disputes whether those models should be screened out or down-weighted in projections. **What's contested (the specific claim).** Hausfather et al.'s prescription that the modeling community should reject/down-weight CMIP6 models whose sensitivity falls outside the IPCC AR6 "likely" range. The counter-claim: excluding models by their *global* sensitivity is invalid for *regional* impact studies because there's no reliable correlation between a model's sensitivity and its skill on regional impact drivers — so blanket screening discards physically plausible, policy-relevant tail outcomes. A named-procedure dispute, not "models are hard." - *Pro-screening:* Zeke Hausfather, Kate Marvel, Gavin Schmidt (NASA GISS), John Nielsen-Gammon, Mark Zelinka. Stake: keeping impact/economic projections from being biased warm. - *Anti-blanket-exclusion:* Ranjini Swaminathan, Richard Betts, Chris D. Jones, Jacob Schewe, Christopher Reyer et al. (Reading/NCEO/Met Office/PIK). Stake: not under-sampling high-impact tail risk in adaptation planning. A separate McDonnell et al. group finds the skill benefit of discounting is mixed. - **[double-checked]** Hausfather, Marvel, Schmidt, Nielsen-Gammon, Zelinka, *"Climate simulations: recognize the 'hot model' problem,"* Nature 605, 26–29 (2022), DOI 10.1038/d41586-022-01192-2, https://www.nature.com/articles/d41586-022-01192-2 (PMID 35508771) — argues the community should weight models by how well they reproduce other evidence and stop treating all models as equally likely. The loud claim. - **[double-checked]** Swaminathan, Schewe, Jones, Betts, Reyer et al., *"Regional Impacts Poorly Constrained by Climate Sensitivity,"* Earth's Future 12, e2024EF004901 (5 Dec 2024), DOI 10.1029/2024EF004901, https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2024EF004901 — as surfaced: "there is no universal relationship between EffCS and projected changes in a number of important climatic drivers of regional impacts," so "excluding high-sensitivity models may not be necessary for regional impact assessments." Directly rebuts the screening prescription. Press: University of Reading, *"Impact studies should include high-sensitivity climate models,"* https://www.reading.ac.uk/news/2024/Research-News/Impact-studies-should-include-high-sensitivity-climate-models . - Complicating: McDonnell et al. (2024), *"To What Extent Does Discounting 'Hot' Climate Models Improve … Predictive Skill?"* Earth's Future, DOI 10.1029/2024EF004844 — helps for some variables, mixed regionally. **How long unresolved.** Live since the May 2022 Nature comment; still argued in Dec 2024 Earth's Future papers. ~3+ years. **Mappable from outside?** Yes — the loud claim, the down-weighting method (IPCC AR6), and the rebuttals are all published and open. CMIP6 model output is public. **Confidence genuinely unresolved: HIGH.** Clear two-sided, named-party, published disagreement, with the rebuttal postdating the claim by 2.5 years. ---
**The area.** Influential modern minimum-wage studies using "bunching" and stacked event-study designs report little to no job loss; critics say the near-zero result is an artifact of design choices. **What's contested (the specific claim).** Cengiz, Dube, Lindner & Zipperer (2019) and Dube & Lindner (2024) find own-wage employment effects near zero (Dube–Zipperer meta median own-wage elasticity ≈ −0.13). Neumark & Rodriguez-Lopez claim that, with the same data and event-study architecture, specific researcher choices drive the null — notably (a) a bunching window that assumes effects are exactly zero outside a narrow wage band (missing job losses across the wider distribution) and (b) measuring employment as a share of a population that itself falls with the minimum wage. Reverse those choices and job loss returns. - *"Modest/no effect":* Arindrajit Dube, Attila Lindner, Ben Zipperer, Doruk Cengiz. Stake: the dominant modern-methods reading used in policy debates (federal $15, CBO scoring). - *"Still reduces employment":* David Neumark, Antonio Rodriguez-Lopez (UC Irvine); Neumark & Shirley. Stake: defending the traditional finding; heavily politicized on both sides. - Cengiz, Dube, Lindner, Zipperer, *"The Effect of Minimum Wages on Low-Wage Jobs: … a Bunching Estimator,"* NBER WP 25434 / QJE (2019), https://www.nber.org/papers/w25434 . The loud near-zero result. - **[double-checked]** Neumark & Rodriguez-Lopez, *"Modern Difference-in-Differences, Same Old Answer: What Explains the Small Estimated Effects of Minimum Wages on Employment?"* (working paper, 2024/25), https://sites.socsci.uci.edu/~jantonio/Papers/ModernDiD_SameOldAnswer.pdf — as surfaced: the null results are "artifacts of specific researcher choices that, once corrected, yield the same old answer: minimum wage increases reduce employment," and Dube's findings "are fragile and depend critically on a number of choices regarding variables, events, sample definitions, and weighting." Directly states the dispute. (The two specific mechanisms are paraphrased from search summaries — verify against the PDF text.) - Dube & Zipperer, *"Own-Wage Elasticity …,"* NBER WP 32925 (2024), https://www.nber.org/papers/w32925 (restates the modest-effect meta finding); Neumark & Shirley, *"Myth or Measurement …,"* NBER WP 28388 (2022), https://www.nber.org/papers/w28388 . **How long unresolved.** The methods fight runs from 2019; the Neumark–Rodriguez-Lopez rebuttal and Dube–Lindner Handbook chapter are 2024, with Dube–Zipperer 2024 restating the other side. Active in 2024–2025. **Mappable from outside?** Yes — all papers public; replication runs on shared public data (CPS/QCEW), which makes the design-choice disagreement unusually checkable. A strong candidate for actually re-running. **Confidence genuinely unresolved: HIGH** it's live and contested; MEDIUM only on the precise in-text wording of the Neumark–Rodriguez-Lopez mechanisms. ---
**The area.** Companies sell carbon-removal credits for spreading crushed silicate rock on farmland, but the field disputes whether standard measurement methods actually prove durable CO₂ removal or systematically overestimate it. **What's contested (the specific claim).** That cation-based measurement (tracking Ca/Mg released by weathering, converted to CO₂-removal figures) reliably quantifies real removal for crediting. Critics claim cation-based approaches can overestimate — cations may re-precipitate into secondary minerals or exchange, and the inorganic-carbon sink may not materialize — so inorganic-carbon-based measurement is needed to check them. A "loud claim nobody independently checked" (credits sold on model-based estimates) crossing into "nobody's job" between geochemistry, soil science, and carbon-market registries. - *Optimistic/commercial:* ERW developers reporting large drawdown, plus crediting registries writing methodologies. Stake: a fast-growing CDR credit market. - *Cautionary:* Hasemer et al. (measurement-methods critique), Power et al. (rates overestimated), CarbonPlan (independent analyst). Stake: integrity of carbon-removal claims; avoiding paid-for phantom removals. - Hasemer, H., et al., *"Measuring enhanced weathering: inorganic carbon-based approaches may be required to complement cation-based approaches,"* Frontiers in Climate 6 (2024), DOI 10.3389/fclim.2024.1352825, https://www.frontiersin.org/journals/climate/articles/10.3389/fclim.2024.1352825/full — as surfaced: "CDR estimated by cations is often greater than CDR estimated by inorganic carbon," and "the unclear fate of cations could lead to overestimations." The title itself states the disagreement. - Power, I. M., et al. (2024), *"Are enhanced rock weathering rates overestimated? A few geochemical and mineralogical pitfalls,"* summarized at https://carbondioxide-removal.eu/en/2025/01/02/power-et-al-2024-are-enhanced-rock-weathering-rates-overestimated-a-few-geochemical-and-mineralogical-pitfalls/ — title poses the contested claim. - CarbonPlan, *"Does enhanced weathering work? We're still learning,"* https://carbonplan.org/research/enhanced-weathering-fluxes — independent, non-commercial analyst treating net removal as unproven. **How long unresolved.** The measurement/verification dispute intensified 2023–2025 as crediting scaled. Unresolved and active now. **Mappable from outside?** Partly. The scientific measurement disagreement is mappable from public papers. But the strongest "loud claim nobody checked" leg — specific companies' per-project removal numbers and their MRV — is largely proprietary/behind registry methodologies. **Confidence genuinely unresolved: MEDIUM-HIGH.** Clearly contested measurement methodology with multiple 2024 critiques; softer than the top tier because it's more "methods not yet trustworthy" than one headline result with a named defender vs. named rebutter. ---
# Run 0 — candidate areas worth mapping
**Prepared:** 2026-08-09.
**Method in one line:** five parallel domain sweeps (semiconductors/hardware, AI/ML internals, quantum, biology/medicine, and climate/materials/economics), then a second independent corroboration pass on the load-bearing quotes. Twelve full candidates below, two shorter honorable mentions, then a ranking and a blunt notes section.
> **Read this before you trust a single quote.** This environment's network egress proxy blocked *all* direct page fetches — arxiv.org, nature.com, science.org, journal sites, blogs, everything returned a 403 policy denial. Only keyword web search worked. So every quote below is **search-index-confirmed, not page-verified**: the URL resolves in the search index, the authors/years/DOIs cross-check across multiple independent queries, and the quoted sentence appears in the engine's grounded summary of that page — but I could not open the PDF or HTML and read the sentence in place. Treat quotes as "accurate to the abstract/snippet as surfaced," one notch below eyeballed. The five or six most load-bearing ones I re-ran through a second, independent search; those are marked **[double-checked]**. Before you build a mapping run on any candidate, open its two anchor URLs on an unrestricted connection and confirm the sentence sits where I say it does. More on this in the notes.
---
## 1. Google's "AlphaChip" RL chip placement vs. its critics — *semiconductors/EDA*
**The area.** Whether Google's deep-reinforcement-learning method for placing the big blocks (macros) on a chip — published in *Nature* in 2021, later branded AlphaChip — actually produces better layouts than a well-tuned classical algorithm and human designers, or whether the headline result falls apart under a fair, reproducible comparison.
**What's contested (the specific claim).** Not "is ML useful in chip design." The sharp claim in dispute: that AlphaChip's RL beats (a) a properly tuned simulated-annealing baseline and (b) commercial tools / human designers on comparable compute. Google says yes and that critics ran the method wrong (no pre-training, ~20× less compute, not trained to convergence). Critics say the original *Nature* comparison wasn't reproducible and omitted steps needed to check it.
**Who's on each side + stakes.**
- *Defenders:* Anna Goldie, Azalia Mirhoseini, Jeff Dean (Google/DeepMind). Authors of the original paper; AlphaChip is a flagship Google "AI-for-chip-design" result taped out in TPUs — reputational and commercial stake.
- *Critics:* Andrew Kahng and Chung-Kuan Cheng's group (UC San Diego / TILOS), authors of the "Assessment of Reinforcement Learning for Macro Placement" line; and Igor Markov (long-time EDA figure), author of "The False Dawn," who also aired integrity allegations.
**Evidence it's live.**
- **[double-checked]** Igor L. Markov, *"The False Dawn: Reevaluating Google's Reinforcement Learning for Chip Macro Placement,"* arXiv:2306.09633 (2023; published in *Communications of the ACM*, 2024). https://arxiv.org/abs/2306.09633 — abstract, as surfaced: *"Crosschecked data indicate that the integrity of the Nature paper is substantially undermined owing to errors in conduct, analysis and reporting,"* and *"Before publishing, Google rebuffed internal allegations of fraud, which still stand."* Directly states the dispute.
- A. Goldie, A. Mirhoseini, J. Dean, *"That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design,"* arXiv:2411.10053 (Nov 2024). https://arxiv.org/abs/2411.10053 — argues the skepticism is "unfounded" and that Cheng et al. "did no pre-training … used [~20×] less compute, and did not train to convergence." The rebuttal.
- Nature, *"Addendum: A graph placement methodology for fast chip design,"* Nature 634, E10–E11 (2024), DOI 10.1038/s41586-024-08032-5. https://www.nature.com/articles/s41586-024-08032-5 — Nature attached an editor's note (Sept 2023) that "the paper's performance claims had been called into question," then published an Addendum rather than a correction.
- Cheng et al., *"An Updated Assessment of Reinforcement Learning for Macro Placement,"* arXiv:2302.11014, **accepted in IEEE TCAD in Dec 2025** — i.e., the critics got a peer-reviewed venue *after* Google's rebuttal and the Addendum. https://arxiv.org/abs/2302.11014 . Original paper: *Nature* 594, 207 (2021), https://www.nature.com/articles/s41586-021-03544-w .
**How long unresolved.** ~4 years (paper mid-2021; critiques 2022–2023; Addendum + Google rebuttal late 2024; critics' peer-reviewed update Dec 2025). Still moving.
**Mappable from outside?** The most mappable of the hardware candidates. Original code (`circuit_training`) and the TILOS/MacroPlacement benchmark repo are public; both sides' papers are public. What an outsider *cannot* touch: Google's internal pre-training corpus, its production placements, and the exact compute used — so you can adjudicate the reproducible-benchmark question but not the "it works in shipping TPUs" claim.
**Confidence genuinely unresolved: HIGH.** Both sides published in 2024–2025; the Addendum did not make critics stand down; no consensus exists.
---
## 2. Microsoft's topological (Majorana) qubit — *quantum / hardware*
**The area.** Whether Microsoft's InAs–Al hybrid nanowire devices actually host Majorana zero modes and constitute a "topological qubit" (the Feb-2025 "Majorana 1" announcement), or whether the signatures are trivial disorder artifacts.
**What's contested (the specific claim).** Two linked claims: (a) that Microsoft's "topological gap protocol" (TGP) reliably identifies a topological phase without false positives, and (b) that the 2025 single-shot parity-measurement device demonstrates a topological qubit. Critic Henry Legg argues the TGP is ill-defined and that the readout happened in regions that look gapless/disordered; Microsoft calls the objections minor bugs / a straw man.
**Who's on each side + stakes.**
- *Pro:* Microsoft Quantum; spokesperson Chetan Nayak (Technical Fellow). Enormous stake — "Majorana 1" is Microsoft's entire quantum roadmap and a public product announcement; DARPA benchmarking involvement.
- *Con:* Henry F. Legg (University of St Andrews), plus condensed-matter physicists who voiced doubt at the March-2025 APS meeting. Legg is an independent academic (relatively low commercial stake — notable, given this field's "everyone qualified has a stake" pattern).
**Evidence it's live.**
- **[double-checked]** Microsoft Quantum, *"Interferometric single-shot parity measurement in InAs–Al hybrid devices,"* Nature 638, 651–655 (2025), DOI 10.1038/s41586-024-08445-2. https://www.nature.com/articles/s41586-024-08445-2 — Nature's editors appended an editorial note (as surfaced): *"the results in this manuscript do not represent evidence for the presence of Majorana zero modes in the reported devices."* An editorial disclaimer on the published paper is a strong "live" signal.
- **[double-checked]** Henry F. Legg, *"On the robustness of topological gap detection via transport,"* a *Matters Arising* in Nature 654, E22–E26 (2026), DOI 10.1038/s41586-026-10567-8, published ~24 June 2026, **with a Microsoft Reply in the same issue.** https://www.nature.com/articles/s41586-026-10567-8 — Legg's point (as surfaced): the claimed parity readout "occurred in regions of phase space with considerable disorder that appear gapless," and shifting measurement windows flips the TGP's own verdict on the same device region. Coverage: The Register, *"Boffin claims Microsoft's supposed quantum leap does not compute due to 'basic Python errors'"* (24 Jun 2026), https://www.theregister.com/research/2026/06/24/boffin-claims-microsofts-supposed-quantum-leap-does-not-compute-due-to-basic-python-errors/5260489 ; TechXplore, *"Critique challenges Microsoft's quantum computing claims"* (Jun 2026), https://techxplore.com/news/2026-06-microsoft-quantum.html .
- Foundational preprint comments: Legg, arXiv:2502.19560 (Feb 2025), https://arxiv.org/abs/2502.19560 ; and arXiv:2503.08944 (Mar 2025), https://arxiv.org/abs/2503.08944 .
**How long unresolved.** The Majorana-signature dispute in these devices traces to the retracted 2018 Zhang et al. *Nature* paper; the current TGP/qubit round runs 2023–2026, with the most recent formal *Nature* exchange in **June 2026** — two months before this run. Actively live.
**Mappable from outside?** Partly. The TGP logic and released transport data are analyzable from public documents — Legg re-analyzed released data and pointed at specific code behaviors. But the definitive raw device data and fabrication details are Microsoft-held, and the physics question ("is there a Majorana?") needs experiments only Microsoft runs at scale. You can map the *dispute* thoroughly; you cannot settle the *physics*.
**Confidence genuinely unresolved: HIGH.** A *Nature* Matters-Arising + Reply in June 2026 with neither side conceding.
---
## 3. D-Wave's "quantum supremacy on a useful problem" — *quantum*
**The area.** D-Wave claims its Advantage2 annealer performed a "beyond-classical" simulation of spin-glass quench dynamics that no classical computer can match in reasonable time.
**What's contested (the specific claim).** Whether classical tensor-network methods reproduce D-Wave's results on ordinary hardware. Challengers say yes — they matched the physics "on a laptop." D-Wave says no — the classical methods fail on the *hardest* instances and highest-order measurements, so the supremacy claim stands.
**Who's on each side + stakes.**
- *D-Wave Quantum Inc. (NYSE: QBTS):* a **publicly traded company** whose product narrative and stock lean on the claim; it filed an SEC 8-K defending it. Original paper: A. King et al.
- *Challengers:* Joseph Tindall, A. F. Mello, M. Fishman, E. M. Stoudenmire, D. Sels (Flatiron Institute / Simons Foundation + Boston University). Elite computational-physics reputations at stake; *Science* published both sides.
**Evidence it's live.**
- D-Wave original: King et al., *"Beyond-classical computation in quantum simulation,"* Science (2025), DOI 10.1126/science.ado6285, https://www.science.org/doi/10.1126/science.ado6285 ; preprint arXiv:2403.00910.
- **[double-checked]** Tindall et al., *"Dynamics of disordered quantum systems with two- and three-dimensional tensor networks,"* arXiv:2503.05693, published in *Science* (2026, DOI 10.1126/science.adx2728). https://arxiv.org/abs/2503.05693 — abstract, as surfaced: these experiments "were claimed to be beyond the reach of classical computation … we find that state-of-the-art accuracies can be achieved with modest computational resources." Publicized by the Simons Foundation, *"Quantum Dynamics Breakthrough Overturns Claim of 'Quantum Supremacy,'"* 21 May 2026, https://www.simonsfoundation.org/2026/05/21/quantum-dynamics-breakthrough-overturns-claim-of-quantum-supremacy-opens-new-research-directions/ .
- **[double-checked]** D-Wave's rebuttal, *"D-Wave's Quantum Supremacy Result Stands"* (26 May 2026), https://www.dwavequantum.com/company/newsroom/press-release/d-wave-s-quantum-supremacy-result-stands/ , and SEC 8-K, https://www.stocktitan.net/sec-filings/QBTS/8-k-d-wave-quantum-inc-reports-material-event-04687ff8a7dd.html — position (as surfaced): the "overturned" claim is "inaccurate and not supported by the scientific record" because the new work "does not reproduce the full scope" nor "the hardest problem instances." Technical comment: arXiv:2504.06283.
**How long unresolved.** Claim March 2025; classical challenges March 2025 → *Science* publication May 2026; D-Wave counter-rebuttal May 2026. ~15 months and escalating.
**Mappable from outside?** Highly. Everything is public — two *Science* papers, an arXiv comment-and-reply, company press releases, even an SEC filing. One of the cleanest live disputes here.
**Confidence genuinely unresolved: HIGH.**
---
## 4. Gaussian boson sampling quantum advantage (USTC "Jiuzhang") — *quantum*
**The area.** USTC's photonic Jiuzhang machines claim quantum computational advantage via Gaussian boson sampling (GBS) — sampling from distributions claimed classically intractable.
**What's contested (the specific claim).** Whether real, *lossy* GBS experiments actually beat classical computers, or whether photon loss and noise let a classical algorithm match — or even beat — the experiment's own sample quality. Narrowly: a classical sampler that produces samples *closer to the ideal distribution than the noisy experiment itself* would negate the advantage; USTC's newer experiments claim to beat every proposed classical method.
**Who's on each side + stakes.**
- *USTC* (Jian-Wei Pan, Chao-Yang Lu et al.): flagship Chinese quantum-advantage program; national-prestige stake.
- *Classical challengers:* Changhun Oh, Minzhao Liu, Yuri Alexeev, Bill Fefferman, Liang Jiang (Chicago / Argonne / KAIST). Stake: credibility of the "advantage" milestone. A running cat-and-mouse.
**Evidence it's live.**
- Oh, Liu, Alexeev, Fefferman, Jiang, *"Classical algorithm for simulating experimental Gaussian boson sampling,"* Nature Physics 20, 1461–1468 (2024), DOI 10.1038/s41567-024-02535-8, https://www.nature.com/articles/s41567-024-02535-8 ; preprint arXiv:2306.03709 — abstract, as surfaced: *"Our classical sampler can simulate the ideal distribution better than the experiment can, which calls into question the claims of experimental quantum advantage."* States the disagreement.
- USTC escalation: *"Robust quantum computational advantage with programmable 3050-photon Gaussian boson sampling"* (Jiuzhang 4.0), arXiv:2508.09092 (Aug 2025), https://arxiv.org/abs/2508.09092 — abstract, as surfaced: results "outperform all classical spoofing algorithms, particularly the matrix product state (MPS) method … recently proposed to utilise photon loss to reduce the classical simulation complexity of GBS." Answers the challenge.
- Original claim lineage: *"Quantum computational advantage using photons,"* Science (2020), https://www.science.org/doi/10.1126/science.abe8770 .
**How long unresolved.** Running since 2020; loss/noise-exploiting classical attacks 2022–2024; USTC's answer 2025. ~5–6 years, still cycling.
**Mappable from outside?** Highly — the whole back-and-forth is on arXiv and in Nature/Science/Nature Physics. Weak spot in the *science* (not the mapping): raw experimental sample sets aren't always released, so challengers work from published statistics — part of why it never fully closes.
**Confidence genuinely unresolved: MEDIUM-HIGH.** Genuinely ongoing, but the pattern is escalation (each new machine resets the target) rather than a clean two-party fight over one paper.
---
## 5. Do "reasoning" models actually reason, or is the reported "accuracy collapse" a test artifact? — *AI/ML internals*
**The area.** Apple's high-profile claim that frontier "reasoning" models suffer a "complete accuracy collapse" past a complexity threshold on planning puzzles is disputed as an experimental-design artifact — and the rebuttals are themselves disputed.
**What's contested (the specific claim).** Shojaee et al.'s (Apple) finding that models like Claude, o3, and Gemini 2.5 fail *fundamentally* on Tower of Hanoi / River Crossing beyond a complexity threshold. The rebuttal claims the "collapse" is (a) models hitting output-token limits (and saying so), (b) an autograder scoring truncation as reasoning failure, and (c) River-Crossing instances that are mathematically unsolvable for N>5 — so models were penalized for refusing impossible problems. The counter-claim (Marcus) is that the rebuttals "fall short" and the limitation stands.
**Who's on each side + stakes.**
- *Reasoning-skeptic:* P. Shojaee, I. Mirzadeh et al. (Apple). Apple is a frontier-model laggard; a "reasoning is illusory" result is reputationally convenient — critics note this.
- *Rebuttal:* C. Opus & A. Lawsen (Open Philanthropy-affiliated), arguing methodology flaws.
- *Defends Apple / attacks the rebuttals:* Gary Marcus. Long-standing public stake in LLM skepticism.
**Evidence it's live.**
- Apple: Shojaee et al., *"The Illusion of Thinking …,"* arXiv:2506.06941 (June 2025), https://arxiv.org/abs/2506.06941 (also https://machinelearning.apple.com/research/illusion-of-thinking ).
- **[double-checked]** C. Opus & A. Lawsen, *"Comment on 'The Illusion of Thinking' …,"* arXiv:2506.09250 (2025), https://arxiv.org/abs/2506.09250 — abstract, as surfaced: findings "primarily reflect experimental design limitations rather than fundamental reasoning failures … Tower of Hanoi experiments systematically exceed model output token limits at reported failure points, with models explicitly acknowledging these constraints." The direct rebuttal.
- Gary Marcus, *"Seven replies to the viral Apple reasoning paper — and why they fall short,"* June 2025, https://garymarcus.substack.com/p/seven-replies-to-the-viral-apple — the argument continuing *after* the rebuttal. Also *"Rethinking the Illusion of Thinking,"* arXiv:2507.01231 (July 2025). Upstream skeptic line: Apple's *GSM-Symbolic*, arXiv:2410.05229 (ICLR 2025).
**How long unresolved.** GSM-Symbolic Oct 2024; *Illusion of Thinking* June 2025; rebuttal + counter-defense within days; counter-rebuttal July 2025. Active, no concession.
**Mappable from outside?** Yes, largely — and you could partly *re-adjudicate* it: the papers, rebuttal, and counter-defense are public, the puzzles are reproducible, and the token-limit argument is checkable by anyone with API access. Only semi-hidden piece: the exact model versions / decoding settings Apple used.
**Confidence genuinely unresolved: HIGH.**
---
## 6. Are "emergent abilities" of LLMs real capability jumps or a metric artifact? — *AI/ML internals*
**The area.** Whether the sudden capability "jumps" with scale that Wei et al. reported are genuine phase transitions, or an illusion created by discontinuous evaluation metrics.
**What's contested (the specific claim).** Schaeffer et al. claim emergence "appears due to the researcher's choice of metric" — swap exact-match/accuracy for a continuous metric and the sharp jump smooths out, so emergence isn't a fundamental property of scaling. Wei's camp concedes some tasks smooth under continuous metrics but argues exact-match *is* what we care about for many tasks, and some capabilities remain abruptly emergent even under continuous metrics. Live sub-dispute (2025): whether *any* residual emergence survives once you control for metric and random-seed variance.
**Who's on each side + stakes.**
- *"Mirage":* Schaeffer, Miranda, Koyejo (Stanford). Paper won a NeurIPS 2023 outstanding-paper award — reputational investment in the deflationary claim.
- *"Emergence is real":* Jason Wei (then Google) et al. — authors of the claim under attack.
- *Keeping it open:* *Random Scaling of Emergent Capabilities*, arXiv:2502.17356 (Feb 2025), argues continuous *loss* can still be bimodal across seeds — i.e., some emergence persists.
**Evidence it's live.**
- Schaeffer, Miranda, Koyejo, *"Are Emergent Abilities of Large Language Models a Mirage?"* NeurIPS 2023, arXiv:2304.15004, https://arxiv.org/abs/2304.15004 — abstract, as surfaced: emergent abilities "appear due to the researcher's choice of metric rather than due to fundamental changes in model behavior with scale." States the disagreement, against Wei et al.
- Jason Wei, *"Common arguments regarding emergent abilities,"* 2023, https://www.jasonwei.net/blog/common-arguments-regarding-emergent-abilities — concedes some curves smooth under log-prob/edit-distance metrics but argues hard metrics like exact match are what you actually want. Original target: Wei et al., *"Emergent Abilities of Large Language Models,"* TMLR 2022, arXiv:2206.07682.
- Continuation: *Random Scaling of Emergent Capabilities*, arXiv:2502.17356, https://arxiv.org/abs/2502.17356 .
**How long unresolved.** Claim 2022; mirage 2023; still generating adjudicating papers in 2025. ~3 years.
**Mappable from outside?** Fully. A public argument over metrics and public benchmark curves (BIG-Bench, GPT-3 arithmetic); an outsider can re-plot the same curves under different metrics. No proprietary data needed.
**Confidence genuinely unresolved: MEDIUM-HIGH.** Both sides credible and still publishing, but the field is drifting toward a "nuanced middle" — softening, not exploding.
---
## 7. The Aβ*56 / Lesné amyloid data-integrity fight — *biology / neuroscience*
**The area.** Whether a foundational 2006 *Nature* result claiming a specific amyloid-beta oligomer (Aβ*56) directly causes memory impairment — a paper that helped steer two decades of Alzheimer's funding toward amyloid — was built on manipulated images.
**What's contested (the specific claim).** Narrowly: whether the western-blot images in Lesné et al. 2006 are genuine. The journal and nearly all co-authors say they were spliced/duplicated/erased and can't be verified; the first author says the retraction is wrong. It stays live because the accused author still publicly rejects the finding and objected to a *second* retraction in Jan 2025.
**Who's on each side + stakes.**
- *Retraction side:* Nature editors + co-authors (Karen Ashe, the senior author; Koh, Kotilinek, Kayed, Glabe, Gallagher). Stake: credibility of the broader amyloid field.
- *Dissent:* Sylvain Lesné (first author), who disputes the retraction and denies doctoring images. Stake: his career (resigned his tenured UMN post, effective March 2025) and ~20 of his papers under scrutiny.
- *Nobody's-job angle:* the detection came from independent sleuths (Matthew Schrag, Elisabeth Bik), not the funders or the journal.
**Evidence it's live.**
- **[double-checked]** *"Retraction Note: A specific amyloid-β protein assembly in the brain impairs memory,"* Nature (2024), DOI 10.1038/s41586-024-07691-8, https://www.nature.com/articles/s41586-024-07691-8 — the co-authors listed "agreed with the retraction … Sylvain Lesné disagreed with the retraction." The retraction shipped over the first author's explicit objection — that single line is the live dispute. (Confirmed near-verbatim on a second independent search.)
- *"Alzheimer's scientist resigns after university finds 'data integrity concerns' in papers,"* Science/AAAS (2025), https://www.science.org/content/article/alzheimer-s-scientist-resigns-after-university-finds-data-integrity-concerns-papers — Lesné resigned from UMN effective March 2025 after a university investigation; he also objected to a separate Science Signaling retraction (Jan 2025). The argument continued past the 2024 retraction.
- Background: Alzforum, *"Sylvain Lesné, Who Found Aβ*56, Accused of Image Manipulation,"* https://www.alzforum.org/news/community-news/sylvain-lesne-who-found-av56-accused-image-manipulation ; original (retracted) paper: Nature 440, 352 (2006), https://www.nature.com/articles/nature04533 .
**How long unresolved.** Allegations public July 2022; 2006 paper retracted June 2024; author still disputing and further retractions landing through 2025. ~3–4 years, open.
**Mappable from outside?** Yes — richly. PubPeer threads, the retraction notices, *Science*'s reporting, and the ongoing investigations are all public. One of the best-documented cases available.
**Confidence genuinely unresolved: HIGH.** Caveat: some read this as a fraud story that is *resolving against* Lesné rather than a symmetric scientific dispute. It is genuinely unresolved as an *argument* (the author has not conceded), but if you want a two-sides-both-plausible fight, this is more one-sided than, say, D-Wave.
---
## 8. Do surrogate endpoints (PFS, ORR) actually predict survival in oncology? — *medicine*
**The area.** Whether the tumor-shrinkage and progression-delay measures that now underpin most FDA cancer-drug approvals reliably translate into patients living longer or better.
**What's contested (the specific claim).** That treatment effects on surrogate endpoints (progression-free survival, objective response rate) are *validated* predictors of overall-survival benefit strong enough to justify approval/prescribing. Trial-level meta-analyses reach opposing conclusions — some report strong PFS↔OS correlation in specific first-line settings, others find correlations "generally low" across oncology — and the camps have not converged.
**Who's on each side + stakes.**
- *Skeptics:* Vinay Prasad & Chul Kim (2015); Bishal Gyawali. Stake: patient protection; identities built on the critique.
- *Defenders/users:* FDA, NCCN guideline-writers, industry sponsors. Stake: approval speed, revenue, regulatory throughput.
- *Live twist:* Prasad — the field's most cited *critic* of surrogate-based approval — became FDA CBER Director and is now reported defending surrogates from inside the regulator. The disagreement is now partly *within one person's own record*, while external critics (Gyawali) continue.
**Evidence it's live.**
- **[double-checked]** Prasad V, Kim C, Burotto M, Vandross A, *"The Strength of Association Between Surrogate End Points and Survival in Oncology: A Systematic Review of Trial-Level Meta-analyses,"* JAMA Internal Medicine 175(8):1389–1398 (2015), https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2323416 (PMID 26098871) — finds the correlations "generally low" and the evidence "limited." The loud skeptic claim.
- *"Surrogate end points in oncology: the speed–uncertainty trade-off …,"* Nature Reviews Clinical Oncology (2025), DOI 10.1038/s41571-025-01007-z, https://www.nature.com/articles/s41571-025-01007-z — frames surrogate use as an unresolved trade-off, not settled science (2025).
- Continued tension: *"US FDA's Prasad: 'We Will Always Embrace Surrogate Endpoints,'"* Citeline/Pink Sheet (2025) — the pro-surrogate stance in the regulator's own words. **Title-confirmed only** (body paywalled + egress-blocked).
**How long unresolved.** Actively disputed since at least 2015; conflicting meta-analyses and the Prasad-at-FDA reversal keep it hot in 2025. ~10+ years.
**Mappable from outside?** Yes — one of the most publicly documented methodological disputes in medicine (dueling meta-analyses, JAMA/NRCO commentary, FDA statements). The individual patient-level trial datasets that would *settle* the correlations are often proprietary to sponsors — a clean "loud claim, verification thin" feature.
**Confidence genuinely unresolved: HIGH** it's contested and live. **Medium** only on the Prasad "always embrace" quote (title-level).
---
## 9. Does epigenetic-clock "reversal" actually mean rejuvenation? — *biology / longevity*
**The area.** Whether a drop in an epigenetic "aging clock" reading after an intervention (partial reprogramming, plasmapheresis, senolytics) is real biological-age reversal or a measurement artifact that doesn't track actual function.
**What's contested (the specific claim).** The claim — loud in longevity biotech — that observing a lower clock age after an intervention *demonstrates* rejuvenation. Critics argue the clocks can't validate this (the "true biological age" of a reprogrammed cell is unknown, so clock predictions are unfalsifiable), that clock reversal can occur without any age-related functional change, and that a clock may register "improvement" merely by suppressing beneficial repair.
**Who's on each side + stakes.**
- *Skeptics:* Kriukov et al. 2024 (associated with Vadim Gladyshev's group, Harvard); a UC San Diego group arguing clock changes are a symptom, not the cause, of aging.
- *Proponents:* reprogramming/longevity labs and companies, plus consumer "biological age" test firms. Large commercial stake — valuations, trial endpoints, and products all lean on clocks being valid readouts.
- *Nobody's-job angle:* nearly everyone qualified to adjudicate built the clock, sells the test, or runs the reprogramming company.
**Evidence it's live.**
- Kriukov D. et al., *"Epistemic uncertainty challenges aging clock reliability in predicting rejuvenation effects,"* Aging Cell (2024), DOI 10.1111/acel.14283, https://onlinelibrary.wiley.com/doi/10.1111/acel.14283 — as surfaced: aging clocks are used to validate rejuvenation during reprogramming, but "their predictions are unverifiable due to the unknown true biological ages of reprogrammed cells." The disagreement, stated against the field's dominant practice.
- UC San Diego, *"Why Our Biological Clock Ticks"* (2024), https://today.ucsd.edu/story/why-our-biological-clock-ticks-research-reconciles-major-theories-of-aging — as surfaced: institutions and companies "are betting on turning back the epigenetic clock … but our research suggests that this may only be treating a symptom of aging, not the underlying cause."
- Supporting: Aging-US, *"Why Epigenetic Clocks May Fail to Measure Anti-Aging Effects,"* https://www.aging-us.com/news-room/why-epigenetic-clocks-may-fail-to-measure-anti-aging-effects .
**How long unresolved.** Escalated 2022–2024 as partial-reprogramming and clock-based claims proliferated; the peer-reviewed pushback (Kriukov 2024) is recent and unanswered at the field level. Open in 2025.
**Mappable from outside?** Mostly — published critiques + the promotional claims they target are public. Caveat: some of the loudest proponent claims (e.g., non-human-primate reprogramming "signals") aren't peer-reviewed and sit inside companies, so part of the proponent side is proprietary.
**Confidence genuinely unresolved: MEDIUM.** Real, published, unrebutted critique from two independent groups — but a diffuse methodological dispute (many parties, no single comment-and-reply) rather than one crisp fight.
---
## 10. CMIP6 "hot models": should high-sensitivity models be down-weighted? — *climate science*
**The area.** About a fifth of the newest generation of climate models have very high sensitivity (equilibrium warming >5 °C for doubled CO₂) and "run hot" vs. observed warming; the field disputes whether those models should be screened out or down-weighted in projections.
**What's contested (the specific claim).** Hausfather et al.'s prescription that the modeling community should reject/down-weight CMIP6 models whose sensitivity falls outside the IPCC AR6 "likely" range. The counter-claim: excluding models by their *global* sensitivity is invalid for *regional* impact studies because there's no reliable correlation between a model's sensitivity and its skill on regional impact drivers — so blanket screening discards physically plausible, policy-relevant tail outcomes. A named-procedure dispute, not "models are hard."
**Who's on each side + stakes.**
- *Pro-screening:* Zeke Hausfather, Kate Marvel, Gavin Schmidt (NASA GISS), John Nielsen-Gammon, Mark Zelinka. Stake: keeping impact/economic projections from being biased warm.
- *Anti-blanket-exclusion:* Ranjini Swaminathan, Richard Betts, Chris D. Jones, Jacob Schewe, Christopher Reyer et al. (Reading/NCEO/Met Office/PIK). Stake: not under-sampling high-impact tail risk in adaptation planning. A separate McDonnell et al. group finds the skill benefit of discounting is mixed.
**Evidence it's live.**
- **[double-checked]** Hausfather, Marvel, Schmidt, Nielsen-Gammon, Zelinka, *"Climate simulations: recognize the 'hot model' problem,"* Nature 605, 26–29 (2022), DOI 10.1038/d41586-022-01192-2, https://www.nature.com/articles/d41586-022-01192-2 (PMID 35508771) — argues the community should weight models by how well they reproduce other evidence and stop treating all models as equally likely. The loud claim.
- **[double-checked]** Swaminathan, Schewe, Jones, Betts, Reyer et al., *"Regional Impacts Poorly Constrained by Climate Sensitivity,"* Earth's Future 12, e2024EF004901 (5 Dec 2024), DOI 10.1029/2024EF004901, https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2024EF004901 — as surfaced: "there is no universal relationship between EffCS and projected changes in a number of important climatic drivers of regional impacts," so "excluding high-sensitivity models may not be necessary for regional impact assessments." Directly rebuts the screening prescription. Press: University of Reading, *"Impact studies should include high-sensitivity climate models,"* https://www.reading.ac.uk/news/2024/Research-News/Impact-studies-should-include-high-sensitivity-climate-models .
- Complicating: McDonnell et al. (2024), *"To What Extent Does Discounting 'Hot' Climate Models Improve … Predictive Skill?"* Earth's Future, DOI 10.1029/2024EF004844 — helps for some variables, mixed regionally.
**How long unresolved.** Live since the May 2022 Nature comment; still argued in Dec 2024 Earth's Future papers. ~3+ years.
**Mappable from outside?** Yes — the loud claim, the down-weighting method (IPCC AR6), and the rebuttals are all published and open. CMIP6 model output is public.
**Confidence genuinely unresolved: HIGH.** Clear two-sided, named-party, published disagreement, with the rebuttal postdating the claim by 2.5 years.
---
## 11. Minimum wage: does the modern "no disemployment" result survive scrutiny? — *economics*
**The area.** Influential modern minimum-wage studies using "bunching" and stacked event-study designs report little to no job loss; critics say the near-zero result is an artifact of design choices.
**What's contested (the specific claim).** Cengiz, Dube, Lindner & Zipperer (2019) and Dube & Lindner (2024) find own-wage employment effects near zero (Dube–Zipperer meta median own-wage elasticity ≈ −0.13). Neumark & Rodriguez-Lopez claim that, with the same data and event-study architecture, specific researcher choices drive the null — notably (a) a bunching window that assumes effects are exactly zero outside a narrow wage band (missing job losses across the wider distribution) and (b) measuring employment as a share of a population that itself falls with the minimum wage. Reverse those choices and job loss returns.
**Who's on each side + stakes.**
- *"Modest/no effect":* Arindrajit Dube, Attila Lindner, Ben Zipperer, Doruk Cengiz. Stake: the dominant modern-methods reading used in policy debates (federal $15, CBO scoring).
- *"Still reduces employment":* David Neumark, Antonio Rodriguez-Lopez (UC Irvine); Neumark & Shirley. Stake: defending the traditional finding; heavily politicized on both sides.
**Evidence it's live.**
- Cengiz, Dube, Lindner, Zipperer, *"The Effect of Minimum Wages on Low-Wage Jobs: … a Bunching Estimator,"* NBER WP 25434 / QJE (2019), https://www.nber.org/papers/w25434 . The loud near-zero result.
- **[double-checked]** Neumark & Rodriguez-Lopez, *"Modern Difference-in-Differences, Same Old Answer: What Explains the Small Estimated Effects of Minimum Wages on Employment?"* (working paper, 2024/25), https://sites.socsci.uci.edu/~jantonio/Papers/ModernDiD_SameOldAnswer.pdf — as surfaced: the null results are "artifacts of specific researcher choices that, once corrected, yield the same old answer: minimum wage increases reduce employment," and Dube's findings "are fragile and depend critically on a number of choices regarding variables, events, sample definitions, and weighting." Directly states the dispute. (The two specific mechanisms are paraphrased from search summaries — verify against the PDF text.)
- Dube & Zipperer, *"Own-Wage Elasticity …,"* NBER WP 32925 (2024), https://www.nber.org/papers/w32925 (restates the modest-effect meta finding); Neumark & Shirley, *"Myth or Measurement …,"* NBER WP 28388 (2022), https://www.nber.org/papers/w28388 .
**How long unresolved.** The methods fight runs from 2019; the Neumark–Rodriguez-Lopez rebuttal and Dube–Lindner Handbook chapter are 2024, with Dube–Zipperer 2024 restating the other side. Active in 2024–2025.
**Mappable from outside?** Yes — all papers public; replication runs on shared public data (CPS/QCEW), which makes the design-choice disagreement unusually checkable. A strong candidate for actually re-running.
**Confidence genuinely unresolved: HIGH** it's live and contested; MEDIUM only on the precise in-text wording of the Neumark–Rodriguez-Lopez mechanisms.
---
## 12. Enhanced rock weathering: is the claimed net CO₂ removal real and verifiable? — *materials / geochemistry × carbon markets*
**The area.** Companies sell carbon-removal credits for spreading crushed silicate rock on farmland, but the field disputes whether standard measurement methods actually prove durable CO₂ removal or systematically overestimate it.
**What's contested (the specific claim).** That cation-based measurement (tracking Ca/Mg released by weathering, converted to CO₂-removal figures) reliably quantifies real removal for crediting. Critics claim cation-based approaches can overestimate — cations may re-precipitate into secondary minerals or exchange, and the inorganic-carbon sink may not materialize — so inorganic-carbon-based measurement is needed to check them. A "loud claim nobody independently checked" (credits sold on model-based estimates) crossing into "nobody's job" between geochemistry, soil science, and carbon-market registries.
**Who's on each side + stakes.**
- *Optimistic/commercial:* ERW developers reporting large drawdown, plus crediting registries writing methodologies. Stake: a fast-growing CDR credit market.
- *Cautionary:* Hasemer et al. (measurement-methods critique), Power et al. (rates overestimated), CarbonPlan (independent analyst). Stake: integrity of carbon-removal claims; avoiding paid-for phantom removals.
**Evidence it's live.**
- Hasemer, H., et al., *"Measuring enhanced weathering: inorganic carbon-based approaches may be required to complement cation-based approaches,"* Frontiers in Climate 6 (2024), DOI 10.3389/fclim.2024.1352825, https://www.frontiersin.org/journals/climate/articles/10.3389/fclim.2024.1352825/full — as surfaced: "CDR estimated by cations is often greater than CDR estimated by inorganic carbon," and "the unclear fate of cations could lead to overestimations." The title itself states the disagreement.
- Power, I. M., et al. (2024), *"Are enhanced rock weathering rates overestimated? A few geochemical and mineralogical pitfalls,"* summarized at https://carbondioxide-removal.eu/en/2025/01/02/power-et-al-2024-are-enhanced-rock-weathering-rates-overestimated-a-few-geochemical-and-mineralogical-pitfalls/ — title poses the contested claim.
- CarbonPlan, *"Does enhanced weathering work? We're still learning,"* https://carbonplan.org/research/enhanced-weathering-fluxes — independent, non-commercial analyst treating net removal as unproven.
**How long unresolved.** The measurement/verification dispute intensified 2023–2025 as crediting scaled. Unresolved and active now.
**Mappable from outside?** Partly. The scientific measurement disagreement is mappable from public papers. But the strongest "loud claim nobody checked" leg — specific companies' per-project removal numbers and their MRV — is largely proprietary/behind registry methodologies.
**Confidence genuinely unresolved: MEDIUM-HIGH.** Clearly contested measurement methodology with multiple 2024 critiques; softer than the top tier because it's more "methods not yet trustworthy" than one headline result with a named defender vs. named rebutter.
---
## Honorable mentions (real, but thinner or softer — not full entries)
- **Do sparse autoencoders recover a model's "real" features?** *(AI/ML interpretability.)* Anthropic's claim that dictionary-learning SAEs recover genuine monosemantic "features" (the "Golden Gate" feature) is contested by *SAEs Do Not Find Canonical Units of Analysis*, Chanin et al., ICLR 2025, arXiv:2502.04878 (SAEs are "incomplete" via stitching and "not atomic" via meta-SAEs) and *Decomposing the Dark Matter of Sparse Autoencoders*, Engels et al., arXiv:2410.14670. Contested claim: Anthropic, *Scaling Monosemanticity* (2024), https://transformer-circuits.pub/2024/scaling-monosemanticity/ . **Why not a full entry:** collegial and partly converging ("SAEs are imperfect but useful"), and the flagship result rests on Anthropic's non-public SAEs/activations — you can test the *method* on open models but not audit the original. Confidence: MEDIUM.
- **Quasi-static negative capacitance in ferroelectric transistors.** *(Semiconductors, device physics.)* Whether HfO₂ ferroelectrics show a genuine *steady-state* negative capacitance (enabling sub-60 mV/dec transistors) or only transient effects + charge-trapping artifacts. Kittl, Houssa, Afanasiev, Locquet, *"Comment on 'Unveiling the double-well energy landscape …,'"* arXiv:2003.00426 (2020) vs. Hoffmann et al., Nature 565, 464 (2019), DOI 10.1038/s41586-018-0854-z. Still generating 2024–2025 papers. **Why not a full entry:** the sharp head-to-head Comment is from 2020; the ongoing dispute is real but diffuse rather than a fresh exchange. Confidence: MEDIUM.
---
## Ranking — which would make the best first mapping run
**Tier 1 — do one of these first.** Clean two-party disputes, both sides still arguing, and unusually *checkable* from public materials:
1. **D-Wave "quantum supremacy" (#3).** The best first run. Everything is public — two *Science* papers, an arXiv comment-and-reply, company press releases, and an SEC 8-K. A publicly traded company on one side gives it stakes and a paper trail most science fights lack, and the fight is *current* (May 2026). You can lay out exactly what each side claims and where the crux sits (hardest instances / highest-order measurements) without needing hidden data.
2. **Minimum-wage bunching/DiD (#11).** The most *re-runnable* candidate. The data (CPS/QCEW) is public and the whole dispute is about design choices on that shared data — a careful outsider could in principle reproduce both results and locate precisely which choice flips the sign. Rare in this list.
3. **AlphaChip (#1).** Public code and a public benchmark repo, a *Nature* Addendum, and a peer-reviewed critique that landed *after* the rebuttal (Dec 2025). Highly mappable on the reproducible-benchmark axis; only the "works in shipping TPUs" leg is walled off.
4. **CMIP6 hot models (#10).** Textbook comment-and-rebuttal with named parties, public model output, and the rebuttal postdating the claim by 2.5 years. The specific-procedure framing ("should we down-weight by global sensitivity for regional impacts?") is exactly the kind of narrow, decidable question this project wants.
**Tier 2 — strong, do after Tier 1.**
5. **Microsoft Majorana qubit (#2)** — the freshest exchange (Nature Matters-Arising + Reply, June 2026) and a genuinely independent critic, but the deciding data is Microsoft-held, so you can map the dispute richly yet never settle the physics.
6. **LLM reasoning collapse (#5)** — vivid, current, partly re-adjudicable with API access; slight risk it reads as a 2025 news-cycle spat rather than a durable open question.
7. **Surrogate endpoints in oncology (#8)** — huge real-world stakes and a delicious "critic-turned-regulator" twist, but it's a decade-old methodological debate more than a single contested result, so mapping it means mapping a *literature*, not a duel.
8. **Aβ*56 / Lesné amyloid (#7)** — superbly documented, but arguably *resolving against* the first author rather than a symmetric live dispute. Pick it if you want a fraud-forensics run; skip it if you want two-plausible-sides.
**Tier 3 — real but softer; do only if a Tier-1/2 topic pulls you toward them.**
9. **Jiuzhang GBS (#4)** and 10. **Emergent abilities (#6)** — both genuinely open but *converging/escalating* rather than head-to-head; the target keeps moving.
11. **Epigenetic-clock reversal (#9)** and 12. **Enhanced rock weathering (#12)** — the most valuable "unexpected field" picks and both have a strong "nobody's job / everyone has a stake" flavor, but each is a diffuse methods critique with part of the loud side sitting inside private companies.
**Not worth doing as written:** the two honorable mentions (SAE interpretability; negative-capacitance transistors) — real disagreements, but collegial/converging or anchored on an exchange that's five years stale. Map them only if you're already in that neighborhood.
---
## Notes on this run
**Blunt version, for you, not a reader.**
**The one thing that actually constrains this run: I could not open a single source page.** Every direct fetch — arxiv.org, nature.com, science.org, APS, Wiley/AGU, journal sites, news, even Wikipedia — was blocked by this environment's egress proxy with a 403 policy denial. Keyword web search was the only channel that worked. So the entire evidence base here is *search-index-grade*: I know each URL resolves, I can see authors/years/DOIs cross-check across multiple independent queries, and the quoted sentence appears in the engine's grounded summary — but nobody in this run laid eyes on the underlying PDF. I re-ran the ~six most load-bearing quotes through a second, independent search and marked them **[double-checked]**; those I'm confident about. The rest I'd rate "very likely accurate, verify before quoting in public." **Before you commit a night to any candidate, open its two anchor URLs yourself and confirm the sentence is where I say it is.** The two I'd re-verify first: the Microsoft-Majorana Nature editorial note wording (#2) and the exact in-text mechanisms in Neumark–Rodriguez-Lopez (#11) — for that one I only saw the abstract/summary, not the body.
**What I searched.** Five parallel domain sweeps — semiconductors/hardware, AI/ML internals, quantum, biology/medicine, and climate+materials+economics — each instructed to prefer PubPeer / Retraction Watch / arXiv comment threads and v2+ revisions / journal comment-and-reply / expressions of concern, and to *drop anything it couldn't source*. Then a consolidation pass (this note) plus a targeted re-check of the anchor quotes. Rough per-domain query themes are listed under each candidate's provenance above.
**What surprised me.**
- **How much of the good stuff is dated 2026.** Today's date in this environment is August 2026, and several of the sharpest exchanges — Legg's *Nature* Matters-Arising on Majorana (June 2026), the Tindall/D-Wave *Science* publication and D-Wave's rebuttal (May 2026) — are within the last three months. That's a feature (genuinely live) but also a caution: the very freshest items had the fewest corroborating sources, so they lean hardest on the search index.
- **The cleanest disputes have a *company* on one side.** D-Wave (an SEC 8-K!) and AlphaChip are the most mappable precisely because a commercial or reputational actor was forced to respond on the record. Pure academic fights (emergent abilities, GBS) are more honest but mushier — they escalate instead of resolving, so "is it still live" is genuinely ambiguous.
**What I'm unsure about / weak spots.**
- **Two candidates aren't symmetric fights.** #7 (Lesné amyloid) and #8 (surrogate endpoints) are arguably "one side is winning" and "a decade-old literature," respectively, rather than two-credible-parties-still-slugging-it-out. I kept them because they're superbly documented and hit the "nobody's job / everyone has a stake" criterion hard, but flagged the asymmetry in each.
- **Quantum is over-represented (3 of 12).** That reflects where the loud-claim-thin-verification pattern genuinely clusters right now, not a thumb on the scale — but if you want more spread, drop GBS (#4, the softest quantum one).
- **Things I deliberately dropped as *resolved* (don't re-add without checking):** LK-99 room-temp superconductor (failed replications; attributed to Cu₂S impurity), Ranga Dias superconductivity (retracted, misconduct finding), Reinhart-Rogoff (long resolved), Google Sycamore 2019 supremacy (eroded by classical spoofing), the deworming "worm wars" (sharpest exchanges ~2015–2019; looked resolving), IBM's 2023 "quantum utility" (the classical-simulability question was largely settled in classical methods' favor), and Othello-GPT linear-vs-nonlinear world models (reconciled, not contested). Cassava Sciences / simufilam I dropped as *near-resolved against Cassava* (drug failed Phase 3; misconduct findings stacking) rather than live — include only if you want a fraud-unwinding run.
- **One quote is title-only:** the Prasad "We Will Always Embrace Surrogate Endpoints" line (#8) — I confirmed the article title but not the body (paywall + egress block).
- **Future-dated index noise:** in raw AI/ML results I saw a few implausible far-future arXiv IDs and discarded them; none are used above. One ID I corrected during the re-check — Apple's *Illusion of Thinking* is arXiv:2506.06941 (2506.09250 is the rebuttal, not the original).
**Bottom line.** Twelve real candidates, all sourced, spanning every domain you asked for plus longevity, geochemistry/carbon-markets, and labor economics. The four Tier-1 picks (D-Wave, minimum wage, AlphaChip, CMIP6 hot models) are the ones I'd actually start a mapping night on — each is a narrow, decidable question with a public paper trail on both sides. The single biggest caveat remains the verification ceiling: everything here is search-confirmed, not page-read, because the network wouldn't let me open the sources.