We organize our results around the two dependent variables: (1) the share of AI conversations devoted to health (intensity, Fig. 1b) and (2) the distribution of those conversations across intent categories (composition). For each, we present a regression model identifying which of the country-level characteristics predict it. Table 1 reports hierarchical regressions for intensity (N = 93), and Table 2 reports the corresponding models for eight intent shares. Robustness checks using alternative merge methods (Supplementary Tables 5–6), false discovery rate (FDR)-corrected P values (Supplementary Table 7), variance inflation factors (Supplementary Appendix section 3), and influential observation diagnostics (Supplementary Fig. 1) appear in the Supplementary Information. Across these checks, the composition associations are shown to be more robust, surviving FDR correction and showing robust effects across most alternative merge methods. Conversely, the intensity–trust association meets our robustness criteria but is more sensitive to specification. As such, the former composition result is best understood as being supported by strong, and the latter by moderate, evidence.
Table 1 OLS regression predicting health conversation intensity
Table 2 Hierarchical OLS regressions predicting eight health conversation intent shares
Health conversation intensity
Confidence in hospitals shows the strongest bivariate correlation with health conversation intensity (r = −0.41, 95% confidence interval (CI) −0.56 to −0.23, N = 100, P < 0.001; Supplementary Table 4): countries where a smaller share of the population reports confidence in hospitals have a larger share of their AI conversations that are health-related. Further, mobile subscriptions (r = −0.29, 95% CI −0.46 to −0.11, N = 108, P = 0.002) and AI readiness (r = −0.27, 95% CI −0.43 to −0.08, N = 107, P = 0.005) are also negatively associated with intensity, while urban population (r = 0.20, 95% CI 0.01 to 0.37, N = 108, P = 0.039), physicians per 1,000 people (r = 0.25, 95% CI 0.06 to 0.42, N = 103, P = 0.011) and government effectiveness (r = −0.23, 95% CI −0.40 to −0.04, N = 108, P = 0.017) also cross the significance threshold. On the other hand, most development and health system indicators, including GDP, internet users and the WHO Universal Health Coverage (UHC) index, show near-zero correlations with intensity. This pattern suggests that general economic development alone does not predict consumer health AI usage.
Our hierarchical regression model outlined in the ‘Empirical strategy’ section in the Methods explains roughly half the cross-country variation in intensity (R2 = 0.532, adjusted R2 = 0.441; Table 1). Overall, institutional and trust characteristics, not broader development indicators, contribute the most additional explanatory power beyond baseline demographics: block 1 (development and demographics) accounts for an R2 of 0.281, though note that no individual predictor reaches significance under HC3 standard errors. Adding block 2 (health system) increases R2 by 0.109, with government health expenditure as a share of GDP reaching significance in the full model (β = 0.46, 95% CI 0.07 to 0.86, z = 2.28, P = 0.022). Lastly, block 3 (institutional and trust) adds the largest increment over block 1 (ΔR2 = 0.141). Overall, our main finding is that confidence in hospitals is the strongest individual predictor in the full model (β = −0.49, 95% CI −0.89 to −0.10, z = −2.43, P = 0.015; P < 0.001 under ordinary least squares (OLS) standard errors).
The negative association between confidence in hospitals and conversational AI usage for health-related intent suggests that countries with lower institutional trust show higher AI health usage. Both confidence in hospitals and government health expenditure are significant in two of the four merge specifications (Supplementary Table 5), with the direction and approximate magnitude of both coefficients remaining stable across all four methods. Substituting the Wellcome Global Monitor 2020 confidence measure for the 2018 measure yields a coefficient in the same direction (Supplementary Appendix section 10), though it is not statistically significant.
Figure 2 shows the country-level pattern: countries where fewer people express confidence in hospitals (for example, Iraq, Iran, Algeria and Libya) cluster at the high end of AI health usage, while countries with high confidence (for example, India, Singapore, Malaysia and Thailand) cluster at the low end. In practical terms, a one-standard-deviation decrease in confidence in hospitals (roughly 13 percentage points) is associated with approximately a 1 percentage point increase in health conversation intensity, with point estimates ranging from −0.18 to −0.49 across the four merge specifications.
Intent composition
Hospital confidence does not reach significance for any of the eight individual intent shares (Table 2). Rather, intent composition is associated primarily with development and health system characteristics.
The bivariate correlations show a consistent development gradient across intents (Supplementary Table 4). Health information and education and research and academic support, the two largest categories, correlate negatively with most development indicators, with the strongest associations in the range of r = −0.50 to −0.65 (P < 0.001). Symptom questions, fitness and lifestyle, and healthcare navigation display the reverse pattern, with the strongest positive correlations reaching r = 0.46 to 0.75 (P < 0.001). The strongest individual associations are internet users with fitness (r = 0.75, 95% CI 0.66 to 0.83, N = 106, P < 0.001), UHC index with Fitness (r = 0.74, 95% CI 0.64 to 0.82, N = 106, P < 0.001) and AI readiness with healthcare navigation (r = 0.72, 95% CI 0.61 to 0.80, N = 107, P < 0.001). Medical paperwork is the exception, with near-zero correlations with most development indicators (population 65+: r = 0.01, internet users: r = 0.05). This pattern suggests medical paperwork intent is associated with a different set of country characteristics.
The hierarchical regressions are consistent with these findings (Table 2 and Fig. 3). Model fit ranges from adjusted R2 = 0.337 for medical paperwork to 0.760 for healthcare navigation, with block 1 (development and demographics) accounting for the explained variance in six of the eight intents, with GDP (log) and population 65+ as the most consistently significant predictors. Healthcare navigation is the best-predicted intent, with four predictors reaching PFDR < 0.05: GDP (log) (β = 0.66, 95% CI 0.30 to 1.02, z = 3.61, P < 0.001, PFDR < 0.01), population 65+ (β = 0.69, 95% CI 0.39 to 1.00, z = 4.45, P < 0.001, PFDR < 0.001), AI readiness (β = 0.44, 95% CI 0.18 to 0.69, z = 3.34, P < 0.001, PFDR < 0.01) and the quadratic GDP term (β = 0.36, 95% CI 0.11 to 0.61, z = 2.80, P = 0.005, PFDR < 0.05).
In raw terms, for a country with average income, a one-standard-deviation increase in log GDP is associated with a 0.7 percentage point increase in the Healthcare navigation share, roughly a 30% increase relative to the sample mean of 2.3%. The other two robust predictors of this intent are comparable in magnitude: a one-standard-deviation increase in the population aged 65+ (about 8 percentage points) corresponds to a 0.7 percentage point higher share (about 32% of the mean) and a one-standard-deviation increase in AI readiness (about 17 points on the 0–100 score) to a 0.5 percentage point higher share (about 20%). GDP (log) and population 65+ also predict higher symptom questions (β = 0.90, 95% CI 0.27 to 1.53, z = 2.81, P = 0.005, PFDR < 0.05 and β = 0.55, 95% CI 0.16 to 0.94, z = 2.78, P = 0.005, PFDR < 0.05, respectively) and lower research and academic shares (population 65+: β = −0.43, 95% CI −0.75 to −0.10, z = −2.59, P = 0.010, PFDR < 0.05). The picture that emerges from Table 2 is a gradient from broad health literacy queries in lower-income, younger-population countries towards specific, personal health management queries in higher-income, older-population countries. Health information and education, the largest category at 43% of health conversations, is the notable exception to this pattern of clear individual predictors. Block 1 explains half its variance (R2 = 0.510), yet no single predictor reaches significance under HC3 standard errors, suggesting that many correlated development indicators share explanatory power without any one dominating.
Medical paperwork is the one intent where the development gradient breaks down. Block 1 explains the least variance for this intent (R2 = 0.113), and block 2 (Health system) adds the largest increment of any intent (ΔR2 = 0.247). The UHC index is the single largest coefficient in our main analysis (β = 0.91, 95% CI 0.46 to 1.37, z = 3.92, P < 0.001, PFDR < 0.001) and health expenditure per capita is also positively associated in the primary specification (β = 0.33, 95% CI 0.10 to 0.56, z = 2.80, P = 0.005, PFDR < 0.05), though this coefficient does not reach significance under the alternative merge methods.
In practical terms, a one-standard-deviation increase in UHC index (roughly 14 index points) is associated with a 4.0 percentage point increase in the medical paperwork share, over half the sample mean of 7.0%. This pattern is consistent with a difference in where the relative unmet needs of individuals lie, that is, where more structured health systems leave users with higher administrative burden and medical paperwork. On the other hand, trust in doctors/nurses points the other way, where lower trust is associated with more paperwork-related queries (β = −0.41, 95% CI −0.73 to −0.10, z = −2.55, P = 0.011, PFDR > 0.05), but this association is not robust enough to be conclusive.
Health expenditure per capita has a different role for other intents. Countries with higher health spending have a smaller share of health conversations about fitness and lifestyle (β = −0.31, 95% CI −0.45 to −0.16, z = −4.15, P < 0.001, PFDR < 0.001) and symptom questions (β = −0.40, 95% CI −0.74 to −0.06, z = −2.28, P = 0.023, PFDR > 0.05). This negative association is consistent with better-funded primary care being associated with lower use of AI for symptom-checking and lifestyle advice. GDP (log) is also positively associated with fitness and lifestyle (β = 0.53, 95% CI 0.13 to 0.92, z = 2.59, P = 0.010, PFDR < 0.05), which indicates that the positive development gradient and the negative health expenditure association operate simultaneously.
Emotional wellbeing is one of the few intents for which block 3 (institutional and trust) contributes the most incremental variance (ΔR2 = 0.073, compared with 0.015 for block 2). Social protection coverage is its strongest bivariate correlate (r = 0.62, 95% CI 0.49 to 0.73, N = 104, P < 0.001), but the coefficient does not reach significance in the full model. The UHC index is negatively associated (β = −0.55, 95% CI −1.03 to −0.07, z = −2.25, P = 0.025, PFDR > 0.05). These patterns are suggestive but not robust enough to be conclusive. Figure 2 illustrates these two patterns. Figure 2b(i)–(iii) shows the development gradient, where higher GDP is associated with a gradient from broad informational queries to specific navigational ones. Figure 2b(iv)–(vi) shows that non-GDP predictors, including physician density, social protection and population age, track distinct intent categories.