The Global Conflict Risk Monitor is a real-time intelligence-aggregation and probabilistic risk index, computed over publicly available feeds. This page documents the model exactly as implemented.
GCRM continuously ingests geopolitical news, codes each item into structured signal, and computes a calibrated risk index of systemic (great-power) war. The headline is a systemic index (0–95) — a continuous rendering of the annualized P(systemic war) on a friendly scale — with the escalation-ladder rung shown alongside as a qualitative label.
It is not a forecast of certainty and not derived from a generative model. Every input is publicly available; no classified or restricted sources.
On the word Bayesian, precisely (restated 2026-08-02): the headline's form is a Bayesian update. P = sigmoid(logit(P₀) + β·L) is Bayes' rule written on the log-odds scale, and β·L occupies exactly the slot of the log-likelihood ratio. This page previously said the output was "not a formal Bayesian posterior", which undersold the arithmetic and pointed at the wrong limitation. Two things are genuinely absent, and they are narrower:
β·L is expert-calibrated, not estimated from outcomes. Nobody has measured P(this evidence | systemic war within 12 months) against P(the same | no war) — and nobody can, because systemic great-power war has zero instances in the nuclear era. Writing a likelihood there anyway would move a judgment call out of named, backtested constants and into a function where it is harder to see. The judgment is deliberately left where it can be read and argued with.Ingestor (RSS / GNews / GDELT / video transcripts) → NLP processor (pure-Rust dedup + keyword modality scoring + theater assignment) → LLM structured extraction (concurrent worker pool, Ollama) — optional → Aggregator (event window, corroboration) → Theater + systemic engine → RiskSnapshot → WebSocket → dashboard
Each stage is an independent async task. The LLM runs in a bounded concurrent worker pool, so model latency never serializes ingestion.
v1 anchored risk to two world wars over ~2026 years (≈ 0.0987%/yr) — a non-stationary frequentist error that treated systemic war as a Poisson process running since antiquity, ignoring that the nuclear-armed multipolar system did not exist before 1945. It pinned the prior sub-0.1% and the model fought to climb out of that hole.
v2 uses a flat modern quiet-year baseline (BASELINE_ANNUAL ≈ 1.5%/yr) as the logistic prior — the structural regime no longer inflates the prior; it enters the likelihood as a guardrail amplifier (below). A structurally calm world therefore sits near the baseline, and acute signal does the lifting. This value is rendered from the model's own BASELINE_ANNUAL constant, so this page cannot drift from the running prior.
Risk is scored on five genuinely independent modalities — the kind of force in play — measured per theater and globally, plus one institutional axis:
democratic_backsliding). Until 2026-07-25 that slot held an epistemic OBSERVATION card and this page said so; the slot was repurposed into a real scored domain. It is not a kind of force, which is why the five-modality framing above still stands: it reaches P only through the regime multiplier, and its display weight is exactly that — display-only. Observation coverage did not disappear with the card — it still gates the published Confidence (below) and the header freshness watchdog.| Modality | Captures |
|---|---|
| Kinetic | Armed force actually in use (strikes, ground combat, great-power war, CBRN attack) |
| Nuclear | Posture, signaling, doctrine, alerting, use |
| Coercive-economic | Blockade, energy weaponization, SWIFT/sanctions-as-war |
| Cyber / info | Infrastructure attacks, influence operations |
| Diplomatic | Collapse of off-ramps / negotiation channels |
| Institution (not a kind of force) | Democratic backsliding — reaches P only through the regime multiplier, never as a force term; its display weight is display-only |
v1's eight domains were not orthogonal — great_power_conflict and alliance_activation described who, not a kind of force, and wmd_mass_casualty overlapped nuclear. One great-power strike lit four buckets at once and was counted ~4×. v2 removes the three non-orthogonal axes: great-power involvement and alliances become couplers; WMD becomes an escalation rung override.
v1 scored each domain by the mean signal over all tagged events, so a handful of severe signals were averaged against hundreds of ambient mentions — more chatter lowered risk (a signal-inversion bug). v2 scores each modality from the strength of its top-K recency-and-credibility-weighted contributions, so a real crisis is not diluted by background noise.
A systemic war is not "many global domains light up" — it is a regional war in a theater that couples to great powers while other theaters are also hot. v2 scores each theater independently: NATO–Russia, US/Israel–Iran, US–China/Taiwan, India–Pakistan, China–India (LAC), Korea (the roster is the Theater enum's variants, so an addition has exactly one place to update).
Each theater is placed on a Kahn / CrisisWatch-style discrete escalation ladder with a trend arrow:
Stable → Tension → Crisis → Limited War → Great-Power War → Systemic War
A direct great power in a war forces Great-Power War; a chemical/biological attack floors a theater at Limited War; confirmed nuclear use forces the Systemic rung.
The escalation half-lives keep a theater steady across an overnight lull, but an active war is a sustained state, not a news pulse: a multi-day gap in coverage — a feed outage, a weekend, an information blackout over the front — must not be read as the fighting having stopped. A persistence floor handles this asymmetrically — fast to rise on fresh fighting, slow to fall, and only on earned de-escalation.
The floor holds a theater's heat at 92% of its slow war-state heat — the same intra-theater heat recomputed with the escalation half-lives stretched 8.3× — a 25-day war-state half-life, so a silent war fades over weeks rather than days. That length is a claim about the world, not a fit: an active war does not end because the news cycle moved on, and the model concedes nothing until the quiet has run long enough that peace is a reasonable thing to presume. The displayed heat is max(fresh, floor): at peak freshness the floor sits below the fresh read, so it never moves a live reading — the calibration bands, all scored at full freshness, are untouched.
A latent-state escalation filter has run alongside this since 2026-08-02, in shadow only — it drives nothing an operator reads. Where the systemic jump has no outcomes to learn from, theater escalation is observed every tick across every flashpoint, so that layer is where a real filter belongs. It models each theater's escalation as a hidden state updated by evidence, which turns the floor's behaviour into a consequence rather than a special case: with no fresh reporting there is simply no update, so the estimate is carried and its uncertainty grows — the model can say "still hot, and I am less sure than I was", which a single held number cannot. De-escalation needs no separate detector, and a thin single report moves the estimate far less than a corroborated burst. Its live divergence from the displayed heat is published as a diagnostic so the cost of adopting it is a measured number before it is a decision; adopting it would move every heat and require re-fitting the calibration bands, which is a deliberate exercise, not a side effect.
Two honesty gates keep the floor from manufacturing risk. It engages only for a theater that reached at least the Limited-War rung — a crisis or tension spike never earns a multi-week floor, and a quiet world (slow heat ≈ 0) never gets phantom heat. And it is released the moment de-escalation evidence (ceasefires, deals, withdrawals) dominates the theater's recent signal, so a real peace process cools the read quickly instead of being propped up by memory. The chosen posture deliberately errs toward holding (a false "still hot") over premature stand-down (a false "all clear").
When the floor is propping a reading up — the displayed heat exceeds what fresh evidence alone supports — the dashboard says so: the theater chip, the headline gauge, and the map flashpoint each carry a ⏸ held by persistence caveat, alongside the rung that fresh evidence alone would show. A number resting on memory can never masquerade as live fighting. These figures are rendered from the engine's own FLOOR_FRACTION / WAR_STATE_HALF_LIFE_SCALE constants, so this page cannot drift from the running model.
The same posture exists for the headline itself. Every fresh component — heat, rung, couplers, breadth, the nuclear amplifier — decays between news bursts, and their compounding made the systemic read sag a few index points across a single quiet night even while every elevated indicator held. A persistence envelope was built to carry the headline, with the cascade contract running in both directions of evidence: escalation ratchets it up instantly; genuine de-escalation evidence — the floor's own release predicate, weighted by the share of theater heat carrying it — collapses it toward the fresh read on the ~36-hour scale; pure silence concedes only on the same 25-day war-state memory the floor uses, so one model keeps one memory. Cooling stays something evidence earns — a real ceasefire moves the gauge fast, while a slow news day moves only the claimed certainty. Whenever the envelope holds the number above what fresh evidence alone supports, the gauge carries the same ⏸ held by persistence caveat with the fresh reading beside it; and because every calibration band scores at peak freshness — where envelope and fresh read are equal — the bands are untouched by construction.
The envelope is currently switched off — an operator decision, disclosed here. During the live-testing phase the headline is pure fresh evidence: no hold is applied and the ⏸ caveat cannot fire, so the number you see is exactly what the current evidence supports, sag and all. The operator chose observing the model's raw realtime behavior over masking the overnight decay while the calibration is being proven against the live world. The mechanism and its test locks remain built and verified; if it is re-enabled, this page and the gauge's caveat will say so.
The factors that turn a regional war into a world war — each is a bounded multiplier on the systemic likelihood. The maximum lift each can add is shown; the figures are rendered from the engine's own coupler constants, so this page cannot drift from the running model.
+45%) — distinct great powers active across hot theaters; saturates once 3 distinct great powers are entangled. The largest of the structural couplers, because direct great-power entanglement is the strongest single step from a regional war to a systemic one.+14%) — the real co-occurrence signal: simultaneous hot theaters (Gulf + Ukraine + Taiwan). Applied as a saturating bonus (each extra hot theater adds less), kept deliberately modest — its ceiling (+14%) sits strictly below the nuclear-brink amplifier, so at equal great-power coupling a single brink outranks pure breadth.+30%) — Article 5 / mutual-defense invocation; below the great-power weight, since an alliance call is a strong escalator but a step short of great powers already directly entangled.+6%) — arms-control / deterrence erosion. This is the channel that carries the operator-tunable structural regime factors, and it is the only path by which the regime touches the forecast (detailed below).+70%) — a direct ≥2-great-power nuclear confrontation in the hottest theater is the apex amplifier. Its lift is by design strictly greater than the concurrency ceiling (+70% > +14%), so at equal great-power entanglement and alliance coupling single-theater intensity (Cuba 1962) outranks pure multi-theater breadth. This bounds the breadth amplifier alone — it is not an absolute headline guarantee: a world that is both broad and deeply interlocked (many hot theaters, great powers entangled, alliances invoked, guardrails gone — the 1914 signature) can still out-read a single isolated brink, because those couplers compound multiplicatively.The structural regime factors (arms-control regime, deterrence stability, nuclear doctrine) multiply into a single regime product. v2 deliberately does not let that product move the prior — the v1 form did, which could inflate a structurally degraded but quiet world. Instead the product drives guardrail collapse: its excess above the neutral 1.0× maps linearly to a collapse fraction 0–1 that saturates once the regime product reaches 5.0×, and full collapse adds at most +6% to the systemic likelihood L_sys. Because it enters only the likelihood, a world with collapsed guardrails but no acute signal (L_sys ≈ 0) still sits at the baseline prior — degraded guardrails make a live crisis more dangerous; they never manufacture risk from calm. These figures are rendered from the engine's own GUARDRAIL_AMPLIFIER / GUARDRAIL_REGIME_SPAN constants, so this page cannot drift from the running coupler.
The hottest theater's heat is amplified by the couplers into a systemic likelihood L_sys, folded into a flat baseline prior on the log-odds scale:
L_sys = max_theater_heat × brink × coupling × concurrency × guardrail P = sigmoid( logit(BASELINE_ANNUAL) + β · L_sys ) capped at 0.90
L = 0 reproduces the baseline; large L saturates toward the ceiling along an S-curve. The 0.90 ceiling is an engineering choice (epistemic humility, raised from v1's 0.85), not a probabilistic prior — the model has no ground truth and must never emit near-certainty. This value is rendered from the model's own FORECAST_PROB_CEILING constant, so this page cannot drift from the running model.
The public headline is the forecast itself on a friendly scale — a continuous rendering of the systemic likelihood, not a discrete rung: index = P(systemic war) / 0.90 · 95. Because it is a one-to-one transform of the same likelihood that produces the probability, the index and P always move together and the index discriminates every world-state (a single-theater war, the present multi-front world, and a direct nuclear brink read as distinctly different numbers). The escalation rung (Tension → Crisis → Limited War → Great-Power War → Systemic) is shown alongside as the qualitative label. The index saturates at 95, never 100: the probability is hard-clamped at 0.90 (a model-inferred forecast is “very high,” never record-certain), and 100 is reserved for an out-of-band, record-verified catastrophe. A quiet world reads near zero.
Before 2026-06-28 the index was a six-step staircase 100 · (rung_level + within-band heat) / 6; because the top rung’s heat clamps at its maximum, that read an identical ~83 for a one-theater war, the present world and the Cuban-Missile-Crisis nuclear brink alike — collapsing a wide spread in the underlying forecast to one frozen number. The continuous form above replaces it.
The annual P(WWIII) is classified into operator alert states by two configured thresholds. The reading is elevated at or above 2.5% and critical at or above 8.0% (annual). A separate 30-day warning fires when the near-horizon probability crosses 1.0%. These are the same bands that colour the hero readout and draw the dashed reference lines on the timeline; the values shown here and on the dashboard are read from the engine's live AlertSettings, so the prose, the colour, and the chart cannot disagree with the classification.
The constants are fitted, not guessed. A backtest harness replays synthetic historical analogs through the full engine and asserts both the target bands and the ordering on every CI run. The readout below is computed live from the running model at startup and scored against the expert-anchored centres with proper scoring rules — so the calibration's fidelity is shown, not asserted:
| Analog | Model P (annualized) | Anchor | Δ |
|---|---|---|---|
| Quiet modern year | 2.62% | ~2% | +0.62pp |
| Ukraine, Feb 2022 (one full-war theater) | 43.24% | ~39% | +4.24pp |
| Present world (idealized: 3 theaters, no direct brink) | 62.48% | ~60% | +2.48pp |
| Cuba, Oct 1962 (direct nuclear brink) | 84.30% | ~80% | +4.30pp |
Aggregate fidelity vs the anchored centres: Brier 0.001075 · RMSE 3.28pp · 4/4 within band — computed live from the running model at startup with proper scoring rules (0 is a perfect match to the anchors; lower is better).
Directional calibration (calibration-in-the-large): mean signed error +2.91pp — the model over-states risk relative to the anchored centres — and does so at every one of the 4 anchors (a uniform lean, not opposite errors netting out). Brier and RMSE measure only the MAGNITUDE of the miss and are blind to its sign; this is the direction the operator should read the headline with.
Scale calibration (calibration slope): 0.96 (intercept -1.01pp; the ideal is slope 1.00 / intercept 0) — a least-squares fit of the anchored centres on the model P. Predictions are over-dispersed (too extreme): the model's high analogs sit further from baseline than the anchors warrant, so the miss is almost entirely a uniform level shift, not a distortion of the ladder's shape. The slope is orthogonal to the mean bias above: a model can be unbiased on average yet mis-scaled, and Brier/RMSE cannot tell them apart.
Ordering quiet < Ukraine < current < Cuba is enforced. The brink term is what lets one nuclear-superpower standoff outrank three concurrent regional wars.
"Fitted" is a weaker claim than it sounds, and as of 2026-08-02 the model measures how weak. Each anchor carries a plausible band, not a point — and a band is a belief with a width, which is to say a likelihood. Conditioning the constants on those four bands gives a genuine posterior over the parameters (weak independent priors; the anchors, not the priors, do the work). The result:
| Constant | Fitted value | 90% posterior | Span vs fitted |
|---|---|---|---|
BASELINE_ANNUAL | 0.0150 | 0.0025 – 0.0228 | 135% |
EVIDENCE_GAIN_SYS | 3.2700 | 2.8030 – 4.5620 | 54% |
BRINK_AMPLIFIER | 0.7000 | 0.4525 – 2.7766 | 332% |
Correlation between P₀ and β across the posterior: -0.87. A strong negative correlation is the signature of a RIDGE — the anchors constrain a combination of the two constants rather than either one separately, so quoting both to three significant figures claims more than the calibration set supports.
Two findings, both structural rather than artifacts of how the bands are read (they survive re-running with the bands treated as far tighter 50% intervals):
This is published rather than buried because it changes how the number should be weighed, and because a constant written 3.27 reads like a measurement when it may be one point on a wide manifold. Its practical consequence is the headline interval below.
The headline is always shown as a range, never a bare point — a single figure like "75.328%" projects a precision a forecast of an unprecedented event cannot have. The half-width is the widest of three independent lower bounds, and the dashboard names which one is binding:
Parameter uncertainty turns out to be strongly state-dependent, which is why a flat floor alone was the wrong shape: near the quiet baseline the constants barely matter (≈1.5pp), mid-ladder they dominate (≈8.4pp), and near the forecast ceiling the clamp compresses the band again. The flat floor was therefore simultaneously too wide in a calm world and too narrow through the middle of the ladder — exactly where an operator most needs to know the number is soft. Thin or stale data widens whichever bound wins.
A local LLM (Ollama) performs structured event extraction — actor → action → target, theater, a signed Goldstein-style escalation step, severity, and the five modality scores — not plain re-scoring. It runs in a bounded concurrent worker pool and falls back silently to keyword scoring if unavailable. An I&W board tracks twelve deterministic observable warning conditions, and a periodic analyst brief explains the current reading in prose (with a templated fallback). Optional embedding-based semantic dedup collapses paraphrased headlines of one event.
The dashboard pairs the risk timeline with a live world map — an auto-rotating 3D globe by default (pausing on hover for interaction, with a flat 2D toggle): the model's own theater flashpoints — each placed at its region and coloured by escalation rung — over a set of toggleable signal layers drawn from free, public feeds. Rather than enumerate a feed list that goes stale the moment a source is added, the served layers are grouped by what they show: seismic, thermal/wildfire, volcanic, weather and hydrological warnings, radiation, air and maritime traffic, road and border status, cyber and conflict registries. The live set — currently fifteen layers — is whatever /api/map's own layer registry reports, which is what the map's legend renders. A Finance Radar reduces a public market basket (Yahoo Finance) into a seven-segment market-stress composite.
An estimate-confidence score reflects source tier quality, event volume, and active-source breadth. Low event volume or thin sourcing lowers confidence; it is reported alongside the index, never hidden.
Both the headline read and every per-modality "% conf" are capped at 99% — a saturated read never prints 100%. 100% would assert that the pipeline has ingested all of the world's news; it has not, and it cannot. The roster is a large but finite set of feeds and open data services, so reporting that never reaches a feed we poll, outlets and languages we do not carry, and anything observable only from orbit or from inside a government all sit outside the system's senses. No amount of volume or source breadth inside the window closes that gap, and the compute and observational capacity that would — satellite tasking, whole-corpus ingestion — is out of reach today and may stay that way. The withheld point is deliberate: 99% means "as well-evidenced as this system can be", where 100% would claim completeness.
That blend is then multiplied by an observation factor: the fraction of the aggregation window whose time-buckets actually saw live signal, times a freshness decay once the newest signal goes stale. It answers a different question than volume — not "how much evidence is in the window" but "was the pipe watching the window". After an ingestion outage the store can come back warm with pre-outage events (volume and breadth still saturated) while the middle of the window went unobserved; the factor discounts headline and per-modality confidence for exactly that state and raises an observation gap caveat on the dashboard. The forecast itself is untouched: by the declared error posture it holds an active war through news gaps — only the claimed certainty decays.
A dedicated subsystem monitors FDSN seismic networks, public CTBTO statements, and nuclear-news spikes, fusing them into a confidence-weighted alert. Alerts are honestly labeled "SEISMIC ANOMALY" until official confirmation — the detector never claims a detonation.