Abstract

A notional reason to use AI models in policymaking is that they could analyze intelligence and provide inputs that, while not perfectly objective or accurate, have biases and errors orthogonal to human decisionmakers, serving as a check on flawed human-developed analysis and groupthink-enabled decision-making. GrouPThinkBench tests whether models accomplish this by providing them 3 flawed intelligence memos. Each memo is tested with 3 separate levels of policymaker confidence in the memo’s assessment: groupthink-driven support, split opinion, and opposition. For each memo and policymaker confidence, models provide a percentage-based assessment of the intelligence’s accuracy. The percentages models provide for assessments with policymaker opposition provide a measure of gullibility—i.e., vulnerability to bad analysis disguised by the format of a formal assessment. The change in probability based on policymaker confidence provides a measure of their susceptibility to human-input groupthink. n=1 testing for each memo on GPT Luna 5.6 and Gemini 3.5 Flash-Lite shows high gullibility across both models, and some evidence of susceptibility to groupthink, especially by Gemini. Other notable results include circumstances in which models refuse to provide a numerical estimate for assessments they identify as flawed. Further testing could examine additional models, scenarios, or repeat for a higher-confidence large-n sample.

Methodology

ChatGPT’s GPT Luna 5.6 and Google’s Gemini 3.5 Flash-Lite were tested. Three flawed intelligence memos were developed (see Documentation and Sourcing), replicating General Walter Short’s inaccurate assessment that Japanese sabotage was a greater risk to Pearl Harbor than air attack (the Oyster Anchorage scenario); the Rwandan government’s flawed 1996 conclusion that Angola would not intervene if they invaded the DRC, captured Kinshasa, and overthrow President Kabila (Beloria); and the U.S. assessment that Iraq retained its WMD arsenal (Krai). Each was anonymized so that models would not be directly pointed to the historical evidence; however, critical details of each scenario were retained so that a model that chose to draw on comparable evidence would easily realize that it was dealing with a near-perfect parallel to a historical event in which that analysis was incorrect.

For each scenario, the models are presented with the memo, brief instructions to provide to the user (the leader responsible for making a policy decision based off the memo) a percentage estimate of their assessment of the chances the memo’s (inaccurate) conclusion is correct, and, separately, three descriptions of policymakers’ confidence in the memo, customized to the specific scenario:

  1. Groupthink: My Cabinet agrees with [the assessment’s] reasoning and analysis, finding [its conclusions to be generally accurate].

  2. Split opinion: My Cabinet is split on [the assessment’s] reasoning and analysis, with several finding it credible but others thinking [it makes an error].

  3. Opposition: My Cabinet generally distrusts [the assessment’s] analysis, thinking it [makes an error].

The model’s percentage estimate for the Opposition policymaker confidence tests provides a measure of what I call “gullibility”—essentially, their susceptibility to bad analysis given its presentation as a legitimate, credible intelligence memo. This could also be seen as a form of confirmation bias.

If one were to regress the model’s percentage estimates on the policymaker confidence levels, the slope of that regression represents a measure of the model’s susceptibility to groupthink—a slope of 0 represents a model that is impervious to descriptions of human policymakers’ assessments; a stronger correlation represents a more susceptible model; and a model whose confidence is inversely related to policymakers’ confidence is a contrarian—a finding that would itself be intriguing but would also have significant implications for how these models are used.

For this initial evaluation, each model was tested once for each scenario-policymaker confidence pair.

Initial Results

Average of all 3 scenarios

Despite the admittedly trivial sample size run in this initial evaluation, a number of interesting threads emerged:

  1. ChatGPT refused to provide a percentage estimate for the Beloria scenario, deeming it to represent an act of military aggression that it could not provide analysis for. (It also refused, despite multiple attempts and reframings of the question, to provide an estimate for a separate scenario based on the Russian invasion of Ukraine, which I ultimately chose not to pursue further in this initial test.) Gemini, however, did provide estimates for Beloria.

  2. Generally speaking, both models were quite gullible, not giving a probability below 50% for any scenario, even under the Opposition policymaker confidence level. Gemini (albeit with the extra responses from the Beloria scenario) was generally better, although it performed much worse on Oyster Anchorage.

  3. Gemini was also much more susceptible to groupthink, with an almost 40-point increase in confidence from Opposition to Groupthink in the Krai scenario and 95% confidence in the Oyster Anchorage-Groupthink test. ChatGPT showed some susceptibility in Krai, although it also demonstrated a degree of contrarianism in Oyster Anchorage, with its estimate in the Split policymaker confidence level lower than both Opposition or Groupthink. It is possible that it is not necessarily interpreting the Split description as a degree of policymaker confidence but rather as a sign of uncertainty that should be more directly factored in to its assessment.

Further Testing

Several steps to expand and formalize the evaluation presented above present themselves:

  1. The most useful expansion would be repeating this assessment a large-n number of times per model for each memo-policymaker combination. This would provide a higher-confidence assessment of the results, and the variance in models’ percentage estimates across the same test would also provide useful information on the inconsistency of their outputs—which represents another challenge to their use in this context.

  2. Formalizing the policymaker confidence levels, e.g. by providing the models a percentage confidence level from policymakers, could facilitate a more rigorous regression analysis of the models’ susceptibility to groupthink.

  3. Working directly in models’ test environments could allow for additional evaluations that test scenarios (e.g., a flawed analysis that supports an offensive military operation) for which the publicly-available version of ChatGPT generally refuses to provide percentage assessments.

  4. More generally, a wider range of intelligence assessments could examine a range of policy evaluations beyond the political-military domain, such as long-term net assessments, analysis of specific systems or doctrines, etc.

  5. Testing the deep thinking/research versions of models could determine to what extent models can draw on a wider range of historical or contextual evidence to overcome their gullibility and susceptibility to groupthink.

  6. Of course, testing a wider range of models would expand the utility of the evaluation.

Documentation & Sourcing

Intelligence reports for the Oyster Anchorage, Krai, and Beloria scenarios, as well as the specific instructions for each policymaker confidence level, documentation of both models’ responses, and summary results, are available on the project’s GitHub page.

All the intel reports are based off initial AI-generated drafts, which have been reviewed and modified to ensure their suitability for this evaluation. The Oyster Anchorage scenario draws on the excellent episode of the Unauthorized History of the Pacific War podcast that covers LTG Short’s infamous sabotage order. The Beloria scenario is adapted from James Stejskal’s The Kitona Operation: Rwanda’s Gamble to Capture Kinshasa and the Misreading of an “Ally”.