1 · Concept overview

Four research traditions sit under this heading and they rarely cite one another. The first is statistical aggregation — Galton’s ox, Condorcet’s jury theorem, the diversity prediction theorem — which is mathematics about estimators and makes no claim about groups as such. The second is the direct measurement of group performance, which produced the collective-intelligence factor and a decade of argument about whether it exists. The third is markets as information aggregators, which has the largest deployed instances and the thinnest calibration record. The fourth is institutional analysis, where Elinor Ostrom asked how real communities govern shared resources and answered with eight design principles, most of which turn out to be about monitoring and punishment. Distributed cognition sits across all four, arguing that the group plus its instruments is literally the cognitive system and the individual mind is the wrong unit of analysis.

The reframe this brief argues for is simple and it is forced by the sources. Both founding artefacts are misreported. Galton observed the median, not the mean, which is a weaker result about aggregation than the retelling. Hardin described an open-access regime, not a commons, and said so himself. And the most-cited modern measurement of group intelligence had a software error in its confirmatory factor analysis whose correction cut the headline variance statistic by more than half. What survives all three corrections is not a story about crowds being wise. It is a design discipline: specific elicitation protocols, specific aggregation rules and specific enforcement machinery, each with a measured effect size and a known way of failing.

This brief also declines to draw dependency edges on the Institute’s technology map, and section 13 explains why at length. Collective intelligence is not a technology that other subjects wait on. It is a property that other subjects on this map exhibit. And it shares a structural problem with the machine end of the same question: the quantity that determines whether an aggregate can beat its best member is error correlation between members, and almost nobody measures it — in silicon or in people.

2 · Current scientific position

Established The founding demonstration is a median, and the difference is not pedantic. At a Plymouth country fair, 800 participants estimated the weight of a slaughtered ox; the median guess was 1,207 pounds against an actual 1,198, within about one per cent, and the source states explicitly that Galton observed the median, not the mean. Frontier The popular retelling substitutes the average, and the substitution silently upgrades the claim, because a median discards the wild outliers a fairground crowd certainly contained. Speculative It is also often said that Galton set out to prove the crowd unreliable; the source consulted here does not characterise his intent, and this brief does not assert it.

Established Condorcet’s jury theorem has two branches and only one is ever quoted. Each voter has an independent probability p of answering a binary question correctly; where p exceeds one half, majority accuracy rises with group size toward certainty; below the threshold the theorem reverses: in the source’s own words, adding more voters makes things worse, and the optimal jury consists of a single voter. Established Berend and Paroush (1998) extended it to heterogeneous competence, including the miracle-of-aggregation case where a small competent minority carries an incompetent electorate. Frontier Romaniega Sancho (2022) found that absent evidence about competence, with individual probabilities following an unbiased distribution, the theorem’s predictions will not hold almost surely. Established Crowd size amplifies whatever the average member is; it does not manufacture competence.

Established The cleanest result in the field is an algebraic identity, not an empirical finding. The diversity prediction theorem states that the squared error of the collective prediction equals the average squared error minus the predictive diversity, so diversity is a subtracted term in an error decomposition rather than a helpful extra, which means an intervention either raises predictive diversity or does nothing. Frontier Every live dispute in the applied literature reduces to whether a given intervention moves that quantity, which is measured on the predictions and not on the people. Speculative The contested relative of this identity, that diversity can outperform ability, has a formal statement and a published mathematical critique; this brief retrieved neither document, so the dispute is recorded as existing and deliberately not characterised.

Established The measured failure modes are specific and none is exotic. The record lists cognitive biases affecting all participants uniformly, social influence degrading accuracy, insufficient diversity producing arbitrary conclusions, consensus-seeking — ETH Zurich experiments where groups attempting to reach a consensus frequently caused accuracy to decrease — and strategic rather than sincere reporting. Established Ariely and colleagues found that averaging repeated estimates from one individual did not significantly improve accuracy, and Rauhut and Lorenz (2011) found that asking oneself more than about three times reduced accuracy below baseline. Frontier The first item sets a ceiling: a bias with zero variance across members contributes nothing to the subtracted term, so no aggregation rule can cancel an error everyone shares.

Frontier Independence is the load-bearing condition, and the literature contains a direct, unresolved contradiction about it. Lorenz, Rauhut, Schweitzer and Helbing, in PNAS in 2011, published the canonical demonstration that social influence undermines the wisdom-of-crowd effect; Ganser and Keuschnigg, in Advances in Complex Systems in 2018, published the opposite, that social influence strengthens crowd wisdom under voting. Established Both records were verified for this brief and neither full text obtained, so this is a claim about the pair rather than about either paper’s internals. Frontier The aggregation rule differs between them — continuous estimate in one case, vote in the other — and asserting that independence is required without first naming the rule is the difference between a true statement and a false one.

Established Information cascades are the formal account of how independence dies, and they are experimentally robust. Bikhchandani, Hirshleifer and Welch (1992) and Banerjee (1992) modelled sequential decisions where prior choices are observable, private signals are not, and every agent updates rationally; a cascade of infinite length can occur on the basis of the decisions of two people. In Anderson and Holt’s urn experiments a cascade appeared in 41 of 56 runs — 73% of trials — including reverse cascades that produced wrong answers when participants prioritised earlier decisions over their own accurate signals. Frontier Cascades are also fragile, reversing when new public information arrives, so what they generate is volatility rather than stable error.

Established The most useful practical result in the subject is a change to the elicitation protocol, and it required no change to the crowd. Asking participants both what they believe is correct and what they expect popular opinion to be, then selecting the answer more popular than predicted, reduced errors by 21.3% versus simple majority votes and by 24.2% versus basic confidence-weighting. Frontier Those figures reached this brief through a fetched encyclopedia article rather than the primary paper, whose publisher was inaccessible, and are printed with that provenance attached. Frontier The significance is categorical: a purely mechanical protocol change beat both standard aggregation rules on the same population, which is a design result rather than an emergence result, and the strongest single reason to treat this subject as engineering.

Established The direct measurement of group intelligence produced a real dataset and a badly damaged headline. Woolley, Chabris, Pentland, Hashmi and Malone studied 192 randomly recruited groups; a collective-intelligence factor explained 43% of variance in the first study and 44% in the second, against 18% and 20% for the next factor. Its predictive validity was the interesting part: the factor predicted criterion tasks including checkers and architectural design, while average member intelligence at r = 0.15 and maximum member intelligence at r = 0.19 did not. Established The correlates were turn-taking equality and social sensitivity, correlated at r = 0.26 and reported as the only statistically significant predictor at b = 0.33, P = 0.05, with the proportion of women largely mediated by it at Sobel z = 1.93, P = 0.03; cohesion, motivation and satisfaction showed no significant correlation at all. Frontier Engel and colleagues replicated in 2014, explaining 49% of between-group variance face-to-face and online. Frontier Then came the hit section 14 treats in full: a software error in the confirmatory factor analysis, corrected in 2022, moved average variance extracted for one-factor models from 44% to 19.6% against a conventional 50% threshold.

Established Prediction markets have the largest deployed instances and the weakest published calibration record of anything here. Berg and colleagues (2008) found that across the five US presidential elections from 1988 to 2004, markets beat 74% of the opinion polls studied, rendered elsewhere as topping the polls 74% of the time across 964 polls; in 2008 the Iowa Electronic Markets predicted the final vote count within half a percentage point. Against a strong benchmark the sign flips: a 2016 randomised experiment found prediction markets performed 12% less accurately than prediction polls using alternative statistical aggregation methods. Established The documented biases are a favourite-longshot bias and Page and Clemen’s finding that prices are pulled toward 50% beyond a year out; the documented failures are Brexit, the 2016 US presidential election, and a 2004 Tradesports manipulation that drove Bush futures to zero before correcting. Established In 2024 over USD 3.3 billion was wagered on the presidential race on Polymarket, and in October a single trader running four accounts with roughly USD 30 million at risk moved Trump’s quoted odds to 53.3% and won USD 85 million. Frontier A market whose price one participant can move that far is not aggregating dispersed information at that moment. It is quoting one person’s belief.

Established Ostrom’s work is the most institutionally detailed thing in the subject, and read carefully it is a result about enforcement. Her eight design principles for enduring common-pool-resource institutions are clearly defined boundaries; congruence between appropriation rules and local labour, material and money; collective-choice arrangements letting those affected modify operational rules; monitoring by auditors who track resource conditions and appropriator behaviour; graduated sanctions; rapid, low-cost, local conflict resolution; minimal external recognition of the right to organise; and nested enterprises for larger systems. Established Governing the Commons appeared from Cambridge University Press in 1990, and Ostrom shared the 2009 Nobel Memorial Prize with Oliver E. Williamson. Count the principles by type: three are enforcement machinery, one exclusion, one constitutional voice, one scale, and only one is about matching rules to the resource. Frontier The commons literature’s actual finding is that collective intelligence at institutional scale is overwhelmingly a monitoring-and-sanctioning problem, which is a very different claim from anything in the aggregation tradition.

Established Hardin, the standard authority for the opposite conclusion, disowned the framing. He said he should have titled his work “The Tragedy of the Unregulated Commons”; what he described was open access, and English common land rights were restricted by law precisely in order to prevent overgrazing and functioned for centuries. Frontier That matters more here than in environmental economics, because the emergentist position rests on an intuition that unmanaged aggregation works. Established Ostrom’s institutional analysis and development framework deflates it further, naming seven rule types — position, boundary, choice, aggregation, information, pay-off and scope — of which aggregation and information are two. Frontier The whole wisdom-of-crowds problem sits inside the institutional framework as two design variables out of seven, which is roughly the weight the measured record supports.

Established Structured elicitation beats unstructured groups, and selected trained forecasters beat everything measured against them. The Delphi method — developed at RAND on a 1944 commission from General H. H. Arnold by Olaf Helmer, Norman Dalkey and Nicholas Rescher — runs anonymised iterated rounds with reasons fed back, and in one documented case predicted new-product sales with 3–4% error against 10–15% for quantitative methods and about 20% for unstructured forecasting. Established The same source’s verdict is that the track record is mixed, that many cases produced poor results, and that if panelists are misinformed the method may only add confidence to their ignorance. The Good Judgment Project, in IARPA’s ACE tournament from 2011 on roughly 100 to 150 geopolitical questions a year, won both seasons, was 35% to 72% more accurate than any other research team on Brier score, and enlisted 260 superforecasters in its final season. Frontier The repeated claim that top forecasters beat intelligence officers with classified access by 30% is carried in the source as a reported figure with broadcast provenance, and is treated here as a reported claim rather than a published statistic.

Established The systems tradition treats the group plus its instruments as the cognitive unit, and it has the best worked example. Hutchins, in Cognition in the Wild (MIT Press, 1995), argued that mental representations are actually distributed in sociocultural systems, and analysed navigation aboard the USS Palau, where pelorus operators, bearing takers, plotters and the captain collectively transform information through a sequence of representational states. Frontier No individual aboard holds the navigation capability; the crew-plus-instruments system does, which makes this the cleanest existing instance of a designed collective outperforming any of its members. Established Stigmergy names the complementary mechanism, indirect coordination through the environment, defined by Pierre-Paul Grassé in 1959 from termite construction. Established The fetched account of its human applications — wikis, open source, Occupy, the 2014 Hong Kong Umbrella Movement — states plainly that it contains no empirical results or metrics for any of them. Frontier Cite the mechanism; do not cite an effect size.

3 · Frontier questions

Frontier The field has just acquired the standard marker of institutionalisation, which is a journal named after itself. Collective Intelligence now publishes directly on the aggregation question, including Douven, Kriegeskorte and Stinson on the stages of crowd wisdom, and Richet and colleagues take up complexity in online collective assessments in Technological Forecasting and Social Change. Established Both are bibliographic records verified for this brief without access to their texts, so they are cited as evidence that the programme is active rather than as evidence for any particular finding.

Frontier The independence question has flipped sign in the peer-reviewed literature and there is no object that reconciles it. What is missing is a classification of aggregation rules by their sensitivity to correlated error: for each rule, the degradation in accuracy as inter-member error correlation rises from zero to one, plotted as a single family of curves. Speculative No such classification exists. Speculative Constructing it is a simulation study plus a modest empirical calibration — unusually cheap for the amount of confusion it would remove, and unclaimed.

Frontier Language models are being used to synthesise populations, and the honest reading of the best result is that most of the signal is demographic. Park, Zou, Kamphorst and colleagues (arXiv:2411.10109, November 2024) built agents from two-hour semi-structured interviews with 1,052 Americans. Established Normalised against human test-retest consistency, the agents reproduced their subjects at 83% from interviews alone, 82% from surveys alone and 86% combined, against a demographics-only baseline of 74%. Frontier The gap that the interview data buys over knowing someone’s demographics is therefore about twelve points on a normalised scale, which is a real result and a much smaller one than the circulating summary. Speculative A synthetic population is a legitimate instrument for power analysis and instrument design, and an illegitimate one for estimating a real distribution, because it has no independent access to the world. Established And a thousand instances of one model is not a crowd in Condorcet’s sense at all: their errors are correlated by construction, so the effective sample size is one.

Frontier Machine forecasters are approaching the aggregate they were built to imitate. Halawi, Zhang, Yueh-Han and Steinhardt (arXiv:2402.18563, February 2024) describe a retrieval-augmented system that searches, forecasts and aggregates, tested on questions published after the underlying models’ training cutoffs, and report that on average the system nears the crowd aggregate of competitive forecasters and in some settings surpasses it. Established No Brier scores were obtainable for this brief, so no number is printed for that result. Speculative If it holds, the substrate changes: the cheapest way to get a crowd stops being people, and the independence condition becomes something a system architect has to engineer deliberately rather than something a population supplies for free.

Frontier Real-time human swarming reports large clinical effects and has attracted almost no adversarial scrutiny. A Stanford study in 2018 reported that radiologists using real-time swarming showed a 33% reduction in diagnostic errors compared with traditional human methods and a 22% improvement over traditional machine learning on chest X-rays; a UCSF study in 2021 reported a 23% increase in diagnostic accuracy over majority voting on MRI images. Frontier The source carrying both numbers gives no sample sizes and contains no criticism of human-swarming claims at all, which is the characteristic shape of an under-scrutinised literature with a commercial sponsor. Speculative These are the largest reported effect sizes in the subject and they are also the least replicated. Both facts should be carried together.

Frontier Prediction markets are being legally reconstructed while their epistemic claims go unaudited. A federal appeals court ruled for Kalshi in October 2024 on a narrow reading of “gaming,” the CFTC dropped its appeal, and a wave of state gaming-regulator suits followed in 2025 and 2026 from Massachusetts, Michigan, Arizona, Minnesota, Nevada, Ohio, Washington and Wisconsin alleging unlicensed gambling. Established Between USD 3 and 5 billion was wagered on NFL games in 2025, which is where the liquidity actually is. Frontier Meanwhile scholars have challenged whether these platforms efficiently and accurately aggregate information about outcomes, and neither of the two dominant venues has a published calibration study in any source consulted for this brief. Speculative For a technology whose entire claim is epistemic, that absence is the story rather than a gap in the citation apparatus.

Speculative The intellectual move the field needs has already been made and is under-adopted. Vermeule’s “Collective Wisdom and Institutional Design” reframes the question from whether crowds are wise to which institutions convert crowd inputs into good decisions. Speculative Fink’s argument for more inclusive collective intelligence converts a fairness claim into a measurable one, because under the diversity identity exclusion is literally a term in the collective error. Frontier Both are bibliographic records here rather than read texts, and both are named because the reframing they perform is the frontier of the subject even where this brief cannot quote their arguments.

4 · Technological bottlenecks

Established The workback target is a designed collective reliably more capable than its most capable member, on tasks that member could attempt alone, measurably and across task types. The qualifiers carry the difficulty: collectives beat their best member routinely on decomposable tasks — that is what a navigation team is — and the impossible version is beating the best member on a task that does not naturally decompose, as a matter of design rather than luck. Frontier Ten intermediate results stand between here and there, and only two are in hand.

Frontier The ladder, in order. First, a task taxonomy with measured decomposability: the ratio of best achievable performance under an optimal split to best individual performance, obtained by decomposition search rather than by inspection. Speculative Second, a classification of aggregation rules by sensitivity to correlated error. Speculative Third, a measured exchange rate between deliberation and independence — information gained per unit of independence lost, as a function of communication bandwidth. Established Fourth, elicitation protocols that recover private information without communication; this link is partly achieved and the surprisingly-popular result at 21.3% is one measured point on the curve. Frontier Fifth, diversity measured in the space that matters, on the predictions as the identity requires rather than on demographics. Established Sixth, a resolution and scoring layer, demonstrated as a method by the Good Judgment Project at 35% to 72% over rival teams. Speculative Seventh, monitoring and graduated sanctions sufficient that reported beliefs are sincere. Speculative Eighth, integration cost below decomposition benefit on a non-decomposable task. Speculative Ninth, transfer across at least three unrelated task families. Handwave Tenth, the standing institution that assembles all nine.

Frontier The binding link is the second one. The fourth and sixth already have real results and could be deployed tomorrow; the first, third, fifth and seventh are hard but conventional measurement programmes; the eighth is an economics problem with a well-understood shape, being Brooks’s Law in software and Williamson’s transaction-cost problem in firms, which makes collective-intelligence design a form of transaction-cost economics with an epistemic objective function. Established The field currently holds two peer-reviewed papers reaching opposite conclusions about whether social influence helps or hurts, and the difference between them is the aggregation rule. Speculative Without the classification, no designer can say of a new collective whether letting the members talk will help it or ruin it, which means every deployment is an experiment and no result generalises.

Frontier The second observation from the ladder is more uncomfortable than the first. The two achieved links are protocol changes requiring no new institution, no new technology and no further research. Speculative The gap between what is known to work and what is deployed is enormous and it is political rather than epistemic, because scored resolution requires an organisation to keep legible records of having been wrong. Handwave On this reading the binding constraint on collective intelligence in practice is not knowledge but the willingness to be measured, and no amount of further research relaxes it.

5 · Research dependencies

Established This subject depends on measurement discipline more than on any technology. Every proposition in the ladder above is a paired comparison between two procedures on a shared item set, with confidence intervals, a stated attempt count and a resolution date fixed in advance. Frontier The methodological standards developed in Intelligence Measurement are prerequisites here rather than refinements, because most effects in this subject are in the tens of per cent and several of its headline results have already failed under re-analysis. Established That is not a hypothetical risk: the two largest corrections in this brief were both produced by re-analysis rather than by new data.

Frontier A second dependency runs through the institutional literature, and the transport is an inference rather than a citation. Ostrom’s framework was developed on common-pool resources — forests, pastureland, fisheries — and the source consulted here states explicitly that it does not address knowledge commons or digital commons. Established The knowledge-commons literature exists; this brief has no source for it and therefore does not claim the eight principles have been validated at digital scale. Speculative Whether they transport probably turns on graduated sanctions, which require stable identity, which is the exact thing online commons do not have. Frontier Monitoring, by contrast, is unusually cheap in a digital commons, so the two halves of Ostrom’s enforcement machinery move in opposite directions when the setting changes.

Speculative A third dependency is theoretical and unglamorous. The subject needs an account of which tasks decompose, and decomposability is currently assessed by eye. Handwave A search over decompositions, scored against best individual performance, would convert the central question from a matter of opinion into a computation, and nothing prevents someone from starting with a modest task battery and a week of compute.

6 · Required experiments

Established Experiment one, and the one that binds: build the correlated-error curves. Simulate majority vote, mean, median, trimmed mean, confidence-weighted mean, surprisingly-popular and market-scoring aggregation over synthetic populations whose pairwise error correlation is swept from zero to one, holding individual accuracy fixed. Frontier Then calibrate on real data by re-analysing existing crowd datasets for which per-item per-person responses survive, measuring the correlation directly rather than assuming it. Speculative The deliverable is a single figure that tells a designer which aggregation rule tolerates how much correlated error, and it is the cheapest high-value paper in the subject.

Frontier Experiment two: the deliberation exchange rate. Run the same estimation task under graded communication bandwidth — no contact, anonymous numeric feedback, anonymous reasons, identified reasons, free discussion — and measure both the information gained and the independence lost, the latter as inter-member error correlation before and after. Speculative Delphi and information cascades are the two endpoints and nobody has measured the interior; the resulting curve would tell any institution where to set its own mixing ratio.

Established Experiment three: a preregistered replication of the collective-intelligence factor with individual-intelligence controls. The confirmatory studies in MBA groups, online gaming communities and cross-cultural groups did not control for individual intelligence scores, and the corrected average variance extracted sits below the conventional threshold. Frontier A clean dataset with individual controls, adequate incentives to defeat the low-effort responding that was independently alleged, and a preregistered factor model would settle a fifteen-year argument in one study.

Frontier Experiment four: publish a calibration study of a real prediction market. Take every resolved contract on a major venue, bin by quoted probability, plot realised frequency against price, and report the reliability diagram and the Brier decomposition. Established The data are public and the analysis is undergraduate-level. Speculative That nobody appears to have published it for either dominant platform, while both advertise epistemic accuracy, is itself the most quotable finding in this section.

Speculative Experiment five: measure strategic misreporting directly. Elicit the same beliefs from the same members under scored and unscored conditions and compare, which turns Ostrom’s monitoring principle into a number and gives the sincerity assumption underneath every aggregation rule its first empirical test.

7 · Engineering requirements

Established Almost nothing in this subject needs new hardware, and that is the engineering finding. The two demonstrated wins — a second-order elicitation question and a scored resolution layer — are software features that fit inside an existing survey tool and an existing issue tracker respectively. Frontier What they need is resolution infrastructure: questions written so that they resolve unambiguously, a fixed resolution date, a named adjudicator, and an immutable record of who said what beforehand. Speculative That is a build of modest difficulty, and its scarcity is itself evidence for the political rather than technical reading of the bottleneck.

Frontier The hard engineering problem is identity, because graduated sanctions require it. Ostrom’s enforcement machinery presumes that the same person can be sanctioned twice, in escalating steps, and that the community can tell. Established Digital commons make monitoring unusually cheap, since every action is logged by default, and make graduated sanctions unusually hard, since identity is cheap to recreate. Frontier This is why online commons converge on reputation systems and why reputation systems are the attack surface: a sanction that can be escaped by re-registering is not graduated, it is a speed bump. Speculative Any serious digital-commons design therefore inherits an identity problem it did not choose and cannot avoid.

Speculative The third requirement is instrumentation for independence. A deployed collective should log per-member per-item responses so that inter-member error correlation can be computed after the fact, which almost no deployed system does. Handwave A platform reporting its own crowd’s effective sample size — the number of genuinely independent opinions behind a headline aggregate — would be the most informative instrument the field could ship, and the computation behind it is standard. Frontier It would also be commercially unattractive, because for most platforms the honest number would be small.

8 · Adjacent technologies

Frontier The sharpest adjacency is to the machine version of the same question, and the two literatures share a structural defect. Multi-Agent Intelligence Systems asks whether a population of language-model agents outperforms one agent, and the assessment there found two measurements missing: no compute-matched multi-agent-versus-single-agent comparison has been published, and no error-correlation matrix has ever been published for a language-model ensemble. Established The human literature has exactly the same hole in a different guise. Frontier Independence is the condition that makes aggregation work — it is the p-independence in Condorcet, the subtracted term in the diversity identity, the quantity that cascades destroy — and it is almost never measured in a deployed human collective either. Speculative Two fields, one silicon and one social, are each unable to say whether their aggregates beat their best members, and both are blocked on the same unmeasured quantity. Frontier That is a real convergence rather than an analogy, and it means the correlated-error curves proposed above would serve both fields at once.

Established The second adjacency is the hinge to cognitive science. Distributed Cognition carries Hutchins’s claim that a navigation team is literally a cognitive system, which makes this subject a special case of Cognitive Architectures rather than a topic in social psychology. Frontier If that framing is right, the memory, action and decision decomposition used to describe individual architectures applies directly to institutions, and an institution can be debugged the way an architecture is.

Frontier The third set of adjacencies is institutional. Institutional Design, Technology Forecasting, Public Policy Foresight and Future Democracies each deploy a mechanism assessed here — structured elicitation, scored forecasting, deliberation, aggregation — and each would benefit from the same two protocol changes. Speculative Reputation Economies is where the identity problem behind graduated sanctions gets its own treatment.

9 · Institutional requirements

Established The institutional record in this subject is unusually well documented on the regulatory side and unusually thin on the epistemic side. The chronology for prediction markets in the United States runs from a 1993 CFTC no-action letter to the University of Iowa, through self-certification under the Commodity Futures Modernization Act in 2000, HedgeStreet’s designation as the first contract market in 2004 and its acquisition by IG Group in 2007, to Dodd-Frank public-interest review powers in 2010 over contracts involving terrorism, assassination, war or gaming. Established The CFTC found Nadex political contracts contrary to the public interest in 2011 and made the same finding against Kalshi in 2023; the October 2024 appellate ruling reversed that direction on a narrow reading of “gaming.” Frontier Federal permission and state prohibition are now in open conflict across at least eight states.

Frontier Regulators have spent thirty years adjudicating whether these venues are gambling and approximately none adjudicating whether they are accurate. There is no regulatory requirement anywhere that a venue advertising probabilistic accuracy publish a reliability diagram for its resolved contracts. Speculative A disclosure rule of that shape — realised frequency against quoted price, by bin, updated quarterly — would cost a platform a few days of engineering and would be the single highest-leverage epistemic regulation available in this area. Frontier It is also the kind of rule that gets written after a scandal rather than before one, and the 2024 single-trader episode is the sort of thing that eventually produces the scandal.

Frontier Inside organisations the binding institutional requirement is a tolerance for scored records. The Good Judgment Project’s instrument works and is almost never adopted internally, because it produces a durable, legible, personally attributable record of having been wrong. Speculative Ostrom’s principles say the same thing from the other direction: monitoring, graduated sanctions and cheap conflict resolution are three of eight, and an institution that will not monitor cannot sanction and therefore cannot sustain sincere reporting. Handwave An organisation that genuinely wanted to be collectively intelligent would begin by scoring its own past decisions and publishing the distribution, and the number that do this is close to zero.

10 · Ethical & societal considerations

Frontier The manipulation question is no longer hypothetical. A single trader with roughly USD 30 million moved a presidential market’s quoted odds to 53.3% and won USD 85 million, and a 2004 manipulation drove a candidate’s futures to zero before the market corrected. Speculative When a market price is reported by news organisations as a public probability estimate, moving it becomes a way to buy a headline, and the cost of doing so is bounded by the market’s depth rather than by anything about the truth. Frontier That is a cheap information-operations channel and it is not covered by any of the gaming-regulation litigation now under way, which asks whether these venues are gambling rather than whether their prices are trustworthy.

Frontier Inclusion becomes an accuracy argument rather than only a fairness one under the diversity identity. If collective error equals average error minus predictive diversity, excluding a group whose predictions differ raises the collective error by a measurable amount. Speculative That converts a normative claim into an empirical one, and it also constrains it: the diversity that helps is diversity in the prediction space, so an inclusion measure that does not move predictions does not reduce error, however otherwise justified. Frontier The honest version of the argument is therefore stronger and narrower than the version usually made for it.

Speculative Synthetic populations create a specific and near-term temptation. Language-model agents reproducing survey subjects at 86% of human test-retest consistency will be cheaper than any consultation, and the pressure to substitute them will fall hardest on exactly the populations that are expensive to reach. Established Their errors are correlated by construction, so a synthetic consultation cannot deliver the independence that makes aggregation informative. Handwave A public body that replaced a consultation with a simulation would be making a category error dressed as an efficiency, and the ethical objection and the statistical objection turn out to be the same objection stated twice.

11 · Civilizational implications

Frontier Every durable institution is an aggregation mechanism with an implicit theory of its own error. Courts, peer review, double-entry bookkeeping, legislative committees and the common law are all procedures for combining fallible judgements, and none of them was designed with the identity in section 2 in view. Speculative Reading them as estimators — asking of each what error correlation it tolerates and what resolution signal it receives — is a research programme that has barely started and that would apply to institutions several centuries old.

Frontier The ceiling result is the civilizationally interesting one. No aggregation rule can remove an error that every member shares, so a civilization’s collective accuracy is bounded by the diversity of its errors rather than by the competence of its members. Speculative That reframes several familiar anxieties as one measurable quantity. Frontier Widespread reliance on a small number of very similar advisory systems is, on this analysis, a systematic reduction in the subtracted term — a monoculture in the error dimension — and it would degrade collective accuracy even if every individual system were more accurate than the humans it replaced.

Handwave The far end of the subject is a genuinely designed collective. If decomposability could be measured, integration cost driven below the difficulty it saves, and error kept from compounding across the decomposition, then a standing institution reliably more capable than its best member becomes an engineering target rather than an aspiration. Speculative Nothing in physics forbids it, the navigation team is an existence proof for the easy case, and the three obstacles are all economic. Handwave A civilization that solved it would have built the first cognitive system whose capability grows with population rather than merely with population’s best draw.

12 · Timelines

These horizons track the measurement programme rather than any technology, because the instruments in this subject already exist and the missing objects are numbers.

  • 10 yr: Frontier The correlated-error curves and the deliberation exchange rate are both cheap enough that either could appear in any year, and the first group to publish them will set the vocabulary for the next two decades. Frontier Expect the collective-intelligence factor to be either rehabilitated by a preregistered replication with individual-intelligence controls or quietly abandoned. Speculative Expect prediction-market calibration studies to become routine as soon as one regulator or one newspaper asks for a reliability diagram, since the data are public and the analysis is trivial.
  • 25 yr: Speculative Scored resolution layers inside real organisations move from curiosity to compliance expectation in at least one high-consequence sector, most plausibly finance or public health, driven by a failure that a scored record would have caught and that an inquiry can name. Speculative Aggregation design becomes a named specialism with a textbook, and the surprisingly-popular protocol becomes the default in professional elicitation the way randomisation became the default in trials.
  • 50 yr: Speculative Hybrid collectives of people and models are the default deliberative instrument, and the interesting engineering is deliberate decorrelation — running architecturally different systems specifically so that their errors differ. Frontier Whether that works at all depends entirely on the curves from the ten-year horizon, which is why the cheap measurement is also the strategically important one.
  • 100 / 250+ yr: Handwave A standing institution that demonstrably outperforms its best member across unrelated task families, with a published error budget and an audited resolution record, would be the first genuinely designed collective intelligence. Handwave It requires no new physics, it has a partial existence proof in every well-run navigation team, and it has never been built.

13 · Technology tree & dependencies

  • Depends on This brief records no depends-on edges, and the decline is a position rather than an omission. Collective intelligence is not a technology that waits on a result some other brief produces. It is a property that groups, markets and institutions exhibit, and the subjects on this map that exhibit it — forecasting institutions, deliberative bodies, multi-agent systems, scientific peer review — are not downstream of it in any sense a dependency edge means. There is one genuine unmet prerequisite, and section 4 states it plainly: a classification of aggregation rules by their sensitivity to correlated error, which nobody is producing and which determines whether any result in the subject transfers to a new deployment. That is a missing scientific result rather than a missing technology, and recording it as a typed edge would misrepresent an open measurement problem as a supply constraint. The prose is where it belongs.
  • Enables Nothing is recorded as enabled here either, for the mirror-image reason. It is tempting to draw arrows outward to every subject that uses an aggregation mechanism, and it would be wrong: those subjects already deploy the mechanisms, in some cases for centuries, and they do not wait on a result this brief produces. The two protocol changes with real measured effects — second-order elicitation and scored resolution — are available to any of them today at negligible cost, which is the opposite of a dependency. What blocks their adoption is an institution’s willingness to keep legible records of being wrong, and that is a political fact about organisations, not an edge on a technology map.
  • Adjacent The honest relationships here are adjacencies and the brief states them as such. Multi-Agent Intelligence Systems is the machine instance of the identical question and is blocked on the identical unmeasured quantity, which makes it the most substantive neighbour rather than a parent or a child. Distributed Cognition supplies the framing under which a group plus its instruments is literally a cognitive system, and Cognitive Architectures inherits that framing directly. Institutional Design, Technology Forecasting, Public Policy Foresight and Future Democracies each operate a mechanism assessed here. Reputation Economies carries the identity problem that graduated sanctions require. Drawing dependency edges to any of these would put a category error into a machine-readable field, and a map that says nothing is more useful than a map that says something false.

14 · Common misconceptions & speculative claims

Established “The average of the crowd guessed the ox’s weight almost exactly.” It was the median, explicitly: 800 participants, 1,207 pounds against an actual 1,198. Frontier The mean of a fairground crowd contains outliers a median discards, so calling it an average makes the demonstration sound stronger than it is — and the whole popular case for aggregation rests on this one anecdote.

Established “Bigger crowds are wiser.” Condorcet’s theorem says the opposite below the competence threshold: where individual accuracy is worse than chance, adding voters makes the majority worse and the optimal jury consists of a single voter. Frontier Under an unbiased distribution of competences with no evidence about competence available, the theorem’s predictions will not hold almost surely. Established Size is an amplifier with a sign attached, and the sign is set by the average member.

Frontier “Wisdom of crowds works as long as you average lots of independent guesses.” Independence is a requirement of some aggregation rules and not others, and the literature contains two peer-reviewed papers reaching opposite conclusions on precisely this point. Established Naming the aggregation rule before asserting the condition is the difference between a true statement and a false one. Established The do-it-yourself version fails outright: averaging repeated estimates from the same person does not significantly help, and beyond about three iterations it hurts.

Established The big one: “collective intelligence is a measurable general factor like IQ.” The 2010 study reported a factor explaining 43% of variance across 192 groups, 44% in a second study, and it became the standard citation for a group analogue of general intelligence. Established In 2021 Riedl and colleagues, several of them authors of the original, reported that a software error had been discovered in the confirmatory factor analysis; a correction published in May 2022 states that average variance extracted values were corrected from 44% to 19.6% for one-factor models. Average variance extracted is a convergent-validity statistic — the share of variance in the measured indicators that the latent construct accounts for — and the conventional threshold for accepting a construct as adequately convergent is 50%. Established The correction’s own text says the corrected values raise concerns about how well the single c-factor model explains indicator variance. Frontier Below 50% the indicators share more variance with things other than the construct than with the construct itself, which is what it means to say a latent factor is not carrying its own measurements; a move from comfortably above the threshold to well below it, caused by a software error, is not a minor adjustment. Frontier The construct did not go to zero: its authors maintain the major conclusions stand, citing factor loadings and meta-analytic data from 22 studies, and Engel and colleagues independently replicated at 49% of between-group variance in 2014. Frontier Credé and Howardson separately attacked the data quality, noting a maximum individual test score of 39 equalling the maximum averaged team score of 39 and an average of 8 out of 50 on an intelligence test, and arguing that low-effort responding could manufacture a spurious general factor through correlated underperformance. Established The confirmatory studies in MBA groups, gaming communities and cross-cultural groups did not control for individual intelligence scores. Frontier The defensible description is contested and partially rehabilitated. Established It is not established, and a brief citing 43% without 19.6% is citing a number its own authors have corrected.

Established “Prediction markets are the most accurate forecasting method we have” — and the statistic usually offered for it, “markets are right 74% of the time.” The 74% is a pairwise win rate against individual opinion polls across 964 polls: not an accuracy rate, not a calibration statistic, and measured against the weakest available benchmark, since no serious forecaster uses a single poll. Established The rest of the record is 12% worse than prediction polls using good statistical aggregation in a 2016 randomised experiment; prices pulled toward 50% beyond a year out; wrong on Brexit and on 2016; and movable by one trader with USD 30 million. Frontier Markets beat single polls, lose to good statistical aggregation, and are systematically miscalibrated at long horizons. Speculative Their real advantage is being cheap, continuous and always available, which is a genuine and underrated virtue and not the same virtue as accuracy.

Established “Groupthink explains institutional failure.” Janis’s construct is culturally dominant and empirically weak. Established Park (1990) found only 16 empirical studies with only partial support, and reported that no research has shown a significant main effect of cohesiveness on groupthink — cohesion being the theory’s central antecedent. Established McCauley (1989) found structural rather than situational conditions predicted it; Aldag and Fuller (1993) called the model too rigidly staged and deterministic; Kramer (1998) found Kennedy and Johnson sought outside expert advice more than Janis suggested; and Choi and Kim (1999) found groupthink symptoms positively correlated with team performance. Frontier The best-known theory of collective failure has substantially failed to replicate. Speculative Use it as vocabulary, not explanation; what it names may be the default behaviour of any group receiving no scored feedback.

Established “Hardin proved that commons fail” — and its mirror image, “Ostrom showed that communities self-organise.” Hardin said he should have titled the work “The Tragedy of the Unregulated Commons”; what he modelled was open access; and English commons rights were restricted by law precisely in order to prevent overgrazing and worked for centuries. Established And Ostrom showed that communities self-organise when eight specific conditions hold, three of them monitoring, graduated sanctions and cheap conflict resolution. Frontier Both halves point the same way: the result is conditional, the conditions are mostly about enforcement, and collective intelligence is a design problem by default rather than an emergent one.

Established “AI can now simulate a population of people.” The result is 1,052 Americans, each interviewed for two hours, with agent accuracy of 86% normalised to human test-retest consistency against a demographics-only baseline of 74%. Frontier Both the sample size and the normalisation are routinely dropped in the retelling, and the headline figure is not a percentage of anything absolute. Frontier More seriously, a thousand instances of one model violates the independence condition by construction, so as a crowd it is a sample of one however many instances are run.

Speculative “Human swarming produces large clinical gains.” The reported figures — 33% fewer diagnostic errors and 22% over machine learning at Stanford in 2018, 23% over majority voting at UCSF in 2021 — come from a source that gives no sample sizes and contains no criticism of the claims, for a technique with a commercial sponsor. Frontier They may well be real; they are the least adversarially examined numbers in this brief, and the appropriate posture is interest rather than belief.

Handwave And the claim this brief is most sympathetic to, offered as speculation rather than smuggled in as fact: the highest-value intervention available today is a scoring layer, not a new institution. Everything here that works — superforecasting, Delphi’s iterated feedback, markets, surprisingly-popular elicitation — has a resolution mechanism, and everything that fails does not. Speculative What would have to be true is that a large fraction of institutional decisions are resolvable in principle, which is contested for exactly the decisions that matter most; that scoring does not induce Goodhart effects; and that organisations accept legible records of being wrong. Frontier The third condition binds, which is why the cheapest well-evidenced intervention in the subject is the one almost nobody adopts.