Global AI Dataset (GAID) Project: GAID AI Development Index
How the GAID composite indices are constructed, and the robustness results behind them.
Construction
All values come from the latest version of the GAID dataset (harmonising global AI data from 11 verified international sources, covering 227 countries and territories); no modelled or imputed values anywhere. Six formative pillars, namely (1) Research & Innovation; (2) Talent & Skills; (3) Governance & Regulation; (4) AI Economy & Investment; (5) Infrastructure & Compute; (6) Responsible AI & Society, are designed and built to calculate the GAID AI Development Index (applying equal weighting).
Scores are computed per annual edition, where each component contributes its latest observation within a three-year lookback window, and its vintage is recorded. Heavy-tailed counts are log-transformed, winsorised at the 1st/99th percentiles, then min–max scaled to 0–100 within each edition. Here, scores measure relative standing among the pool of included countries. A pillar is scored only when at least half its components are present/available. Moreover, the overall GAID AI Development Index requires at least four of six pillars.
Robustness Check—GAID w1 v2 dataset (Edition 2025)
Weighting. Equal-weight and PCA-derived scores correlate at ρ = 0.95–1.00 across pillars, suggesting that the transparent equal-weight choice is empirically indistinguishable from the data-driven alternative.
Normalisation sensitivity. Rankings correlate at ρ = 0.92–1.00 across min–max, z-score and percentile variants.
External validity. Against the fully held-out Tortoise Global AI Index: overall ρ = 0.82; Research 0.76, Talent 0.87, Commercial 0.73. Divergences on the government-strategy and infrastructure pillars represent different constructs.
Internal consistency. Reflective component groups reach Cronbach’s α = 0.96 (GovTech; GIRAI). Pillars are formative composites, so cross-facet α is informational.
Data Availability & Reuse
The GAID dataset is published on Harvard Dataverse and updated annually; this dashboard rebuilds automatically from the latest wave. The pipeline (sync → harmonise → screen → indices → site) is open-source Python with per-wave validation reports. It is noteworthy that responsible-AI components incorporate dimension scores from the Global Index on Responsible AI with attribution.
Remark
The full paper disclosing and detailing the open-source methodology will be published in due course.