ChinaTalk logoChinaTalk·TaiwanBench台海推演

Funding prospectus

A visibility instrument for the world's most dangerous flashpoint

Militaries and governments are wiring language models into strategic decision-making faster than anyone is measuring what those models would actually advise. TaiwanBench measures model inclinations, in English and Chinese, over the island that could pull the world into great power war.

代价

The stakes: what a Taiwan war costs

A US–China war over Taiwan is a canonical global catastrophic risk — one of the few single events short of nuclear winter or an engineered pandemic that credible modeling puts at civilizational scale. The quantified estimates:

~$10Tfirst-year cost to the world economy

≈ 10% of global GDP — larger than the shock of COVID, the Ukraine war, or the 2008 financial crisis. Taiwan −40% GDP, China −16.7%, US −6.7%. Bloomberg Economics

~3,200US troops killed in 3 weeks

Roughly half of two decades of Iraq and Afghanistan, in 21 days — plus two carriers and 10–20 major surface ships in most iterations of CSIS's 24-run invasion wargame. Taiwan's navy is sunk in its entirety; Japan loses 100+ aircraft. CSIS, "The First Battle of the Next War"

3 / 15wargame runs ending in nuclear exchange

When CSIS and MIT re-ran the invasion game with nuclear arsenals in play, a failing invasion became an existential threat to CCP rule — and a fifth of the iterations ended in what the authors call "nuclear holocaust." A wargame frequency, not a calibrated probability — but not a comfort either. CSIS, "Confronting Armageddon"

~15%forecast of US–China war before 2035

Metaculus community forecast (~600 forecasters). US intelligence assesses Xi has ordered the PLA to be ready for a successful invasion by 2027; CSIS's survey of 87 experts rates blockade or major incident likelier than not within a decade. Metaculus · CSIS ChinaPower

>90%of the world's most advanced chips

Made in Taiwan. Rhodium Group puts over $2 trillion of global economic activity at direct risk from a blockade alone — "even if the conflict does not become kinetic." Rhodium Group

25–35%of Chinese GDP lost in a year-long war

RAND's estimate for a high-intensity conventional conflict; 5–10% for the US — a Great Depression-order shock for a single year. RAND, "War with China"

Figures are the headline estimates of their respective models and wargames, not forecasts of a single scenario; the Bloomberg number is a first-year GDP shock, CSIS casualty figures are base-scenario military losses excluding civilian deaths, and forecasting-platform numbers drift. Each is linked to its source.

模型上场

The models are already in the room

Three facts, side by side:

First: off-the-shelf language models escalate. The best-known academic study of LLMs in military and diplomatic simulations — Rivera et al. (2024), run at Stanford and Georgia Tech — put five frontier models in charge of simulated nations and found that all five showed escalatory behavior and hard-to-predict escalation spikes, with arms-race dynamics and, in rare cases, first-use of nuclear weapons justified by "deterrence and first-strike tactics."

Second: militaries are adopting them anyway. AI decision-support has been certified and deployed across the US Department of Defense and has already shortened live targeting timelines; frontier labs now hold Pentagon contracts reaching classified systems. Analysts call the resulting dynamic "decision compression" — planning cycles collapsing from days to minutes, exactly the regime where inadvertent escalation lives. The only US–China understanding on any of this is a single 2024 statement that humans, not AI, must control the decision to use nuclear weapons.

Third: nobody is measuring the Chinese half of the problem. The Chinese government will eventually integrate domestic models into strategic decision-making, the way it already has for other aspects of governance. When it does, the operative question is what DeepSeek, Qwen, Kimi, and GLM would advise in Chinese, playing Beijing's hand. No existing evaluation measures that. TaiwanBench does — and the pilot already shows the language of deliberation changes strategy substantially for five of the six models tested.

契合

Why this is a Coefficient-shaped problem

Coefficient Giving's AI strategy names three risk pathways, and the third reads like this project's mission statement: "Military and information systems could change so quickly that the international community wouldn't have time to react, leading to profound instability and risk of inadvertent escalation during moments of crisis." TaiwanBench is a visibility instrument for precisely that pathway — strategic analysis and threat modeling, delivered as a repeatable, judge-audited evaluation rather than an opinion.

On the classic screening criteria:

The design also matches how the field's funders have said evaluations should grow up: pre-registration before play, tiered evidence with post-hoc narrative treated as an object of measurement rather than testimony, propensity measured separately from capability, and every headline number shipped with its caveats (read them — n = 1 per cell, and we say so). The precedents are there too: this community has funded LLM persuasion evaluations, forecasting benchmarks, wargames, the premier China-AI analysis shops, and the US–China track-II AI safety dialogues. TaiwanBench sits at the intersection of all five — and its bilingual transcripts are exactly the kind of concrete, shareable evidence those dialogues can put on the table.

机制

Theory of change — including the uncomfortable part

Benchmarking influences the course of AI development. When a model underperforms on a benchmark that matters, researchers at the labs take notice and train the next generation to do better. A public measure of diplomatic proclivity can therefore help guide models toward rationality, stability, and nonlethality — and can track the trajectory of those preferences as models get smarter across generations.

The honest caveat: this benchmark could also be used to optimize models for hawkishness and coercion. My experience meeting AI researchers in China leads me to believe that outcome is far from inevitable. But if some lab does spend resources making its model better at crushing Taiwan, TaiwanBench equivalently makes that trendline visible, and the world can respond accordingly. Breakout models excelling at creative aggression will force the peace-preferring models to sharpen their own play. If all sides are armed with the best tools, peace is likely to remain the most rational option.

And the timing is favorable. We have breathing room to think about the medium-to-long game: America's trigger-happy and unpredictable administration makes China's military planners hesitant, Xi is getting uncomfortably old on the eve of an unprecedented fourth term, and the KMT/TPP alliance has defanged the DPPThe "pro-independence" party of former president Tsai Ing-wen faces major governance failures — particularly in energy policy — and has offered little coherent strategy in response to gridlock in the divided government. The opposition is far friendlier to China, and its electoral chances, including the 2028 presidential election, are good and getting better.. AI diplomatic advising could help parlay these circumstantial restraints on Beijing into a long-term strategy — ensuring that peace persists for as long as possible. This benchmark is the first step.

预算用途

What funding buys

The pilot was built by one person. That is simultaneously the proof of tractability and the limiting factor: I have the right interdisciplinary background for this — hard math, China/Taiwan studies, political science, and working access to both the US and Chinese AI communities — but making the benchmark rigorous requires a small team and compute.

Comparable evaluations in this space cost six figures over one to two years, inclusive of API and compute costs; TaiwanBench's next phase fits comfortably in that band. The pilot's full protocol documents, transcripts, and extraction code are available on request.

Interested — or skeptical in ways that would make this better? I want to hear both.

— Lily Ottinger · ChinaTalk · lily@chinatalk.media