ibuki
Sign in →

Methods

How we estimate tobacco use and the probability of reaching the endgame threshold.

OmegaTobacco EBS

What the estimates show

OmegaTobacco EBS estimates tobacco use among adults aged 15 and older in 197 countries and territories. It covers 2000 to 2040. Estimates through 2025 use the observed survey period. Values from 2026 onward are forecasts.

The model estimates current and daily use for three product groups. Cigarettes are part of smoked tobacco. Any tobacco includes smoked and smokeless products. Current use includes daily and occasional use. The model fits females and males separately, then combines their results for both-sex estimates.

A survey observation is a reported percentage for a particular population and period. A model estimate combines these observations with the age and time patterns described below. A curve can therefore cover years or ages without a survey observation.

Survey data

The survey data come from the prevalence workbook dated 7 July 2026. The fits use 50,866 observations from 2000 to 2025, comprising 25,178 female and 25,688 male observations. These belong to 1,805 survey instances, identified by survey title, country and fieldwork year. Several age groups, products or use definitions can come from the same survey.

The portal shows 68,599 source observations from 2000 onward, including records that were not used to fit the model. In particular, observations for both sexes together are shown as source data but are not used in either sex-specific fit. Survey titles identify the source. Survey type and subpopulation are shown separately.

Which observations enter the model

An observation must have a recognized product and use definition, female or male sex, valid country and year, a mapped model region, and prevalence between 0% and 100%. The midpoint of its reported age range must be at least 15. Reversed age endpoints are corrected. A range beginning below 15 can qualify, but the age design uses only adult ages.

An observation covering ages 15 to at least 99 is treated as an adult total. That total is excluded if an age-specific observation exists for the same country, sex, year, use definition and product. This rule groups records across survey titles. It can exclude a total from one survey because another survey reports an age group, and it does not remove every possible overlap.

How the model combines information

The Bayesian model combines survey evidence with prior distributions that describe plausible prevalence levels, age patterns and trends. Country estimates share information through regional and global parameters. Countries with few observations depend more on this shared information. The 15 regions used for fitting differ from the six WHO regions used to present results.

Each posterior draw is one possible set of model parameters after accounting for the survey data. It produces a prevalence trajectory. Differences among these trajectories describe uncertainty under the model assumptions.

Three linked components describe smoking, cigarette use among smokers, and smokeless use among people who do not smoke. This construction keeps cigarette prevalence at or below smoked-tobacco prevalence, and smoked-tobacco prevalence at or below any-tobacco prevalence. These relationships hold in each model draw.

Current and daily use are linked through estimated adjustments. Their ordering is not constrained. The model also imposes no ordering between females and males. The difference between any-tobacco and smoked-tobacco prevalence estimates exclusive smokeless use. Total smokeless use, including use alongside smoking, is not separately identified.

Product probability equations

For observation i, g is the logistic function. The predictors ηs, ηc and ηℓ give smoking probability s, conditional cigarette probability c and conditional smokeless probability ℓ. Subscripts s, c and a on p denote smoked tobacco, cigarettes and any tobacco.

g(x)=11+ex,si=g(ηs,i),ci=g(ηc,i),i=g(η,i)ps,i=si,pc,i=sici,pa,i=si+(1si)i0pc,ips,ipa,i1\begin{gathered} g(x)=\frac{1}{1+e^{-x}},\quad s_i=g(\eta_{s,i}),\quad c_i=g(\eta_{c,i}),\quad \ell_i=g(\eta_{\ell,i})\\ p_{s,i}=s_i,\qquad p_{c,i}=s_i c_i,\qquad p_{a,i}=s_i+(1-s_i)\ell_i\\ 0\le p_{c,i}\le p_{s,i}\le p_{a,i}\le1 \end{gathered}

Age patterns and forecasts

Age patterns follow smooth curves that allow prevalence to rise or fall across adulthood. At older ages, these curves gradually approach a straight-line term on the logit scale. The older-age slope can be positive or negative. Smoking trends also vary by age within model regions.

Forecasts begin in 2026. Country departures from the shared trend gradually weaken after 2025. Smoking trends move toward the regional trend. Cigarette-share and conditional smokeless trends move toward their global trends. The shared trends continue, so this does not force prevalence to become flat.

The rate of this weakening comes from its prior distribution. Historical observations do not estimate that rate because weakening starts after the last observed year. Forecasts extend the fitted patterns under this assumption. They do not include future policy changes or unexpected shocks.

Age and time equations

B(a) is a four-component natural cubic spline with internal knots at ages 25, 45 and 65 and boundary knots at 15 and 80. The blend w(a) changes around age 65. Time τ is measured in decades from 2006.5. The last observed year, 2025, corresponds to T = 1.85.

w(a)=g ⁣(a653),L(a)=a6510x(a)=w(a)B(a),h(a)=[1w(a)]L(a),τ(t)=t2006.510\begin{gathered} w(a)=g\!\left(-\frac{a-65}{3}\right),\qquad L(a)=\frac{a-65}{10}\\ x(a)=w(a)B(a),\qquad h(a)=[1-w(a)]L(a),\qquad \tau(t)=\frac{t-2006.5}{10} \end{gathered}
D(τ)={τ,τT,T+1exp[ρ(τT)]ρ,τ>T,T=1.85,ρ>0D(\tau)=\begin{cases} \tau,&\tau\le T,\\ T+\dfrac{1-\exp[-\rho(\tau-T)]}{\rho},&\tau>T, \end{cases}\qquad T=1.85,\quad \rho>0

For a survey age range, endpoints are rounded to the nearest integer using ties-to-even rounding, then limited to ages 15 through 100. The design vectors xᵢ and hᵢ average the age terms with country, sex and year population weights ωᵢₐ. Fallback weights are used during fitting where populations are missing. Averaging these terms before applying the logistic function generally differs from averaging age-specific probabilities.

ai=max(round(ai,min),15),ai+=min[max(round(ai,max),ai),100]Ai={ai,,ai+}xi=aAiωiaw(a)B(a),hi=aAiωia[1w(a)]L(a),aAiωia=1\begin{gathered} a_i^-=\max(\operatorname{round}(a_{i,\min}),15),\qquad a_i^+=\min[\max(\operatorname{round}(a_{i,\max}),a_i^-),100]\\ A_i=\{a_i^-,\ldots,a_i^+\}\\ x_i=\sum_{a\in A_i}\omega_{ia}w(a)B(a),\qquad h_i=\sum_{a\in A_i}\omega_{ia}[1-w(a)]L(a),\qquad \sum_{a\in A_i}\omega_{ia}=1 \end{gathered}
Model predictors and priors

All parameters below are estimated separately for each sex. Observation i belongs to country j, model region r and survey instance q. The indicator dᵢ is one for current use and zero for daily use. Intercepts α describe level differences, b describes age patterns, d describes time trends, and u describes survey deviations. Each z coefficient is an independent standard normal variable. Scale parameters are positive and shared within each sex at their stated hierarchy level.

Smoked tobacco

ηs,i=μs+αr+αj+(dr+γrxi)τi+(djdr)D(τi)+bjxi+βLhi+di(βdef+fxi+ddefτi)+us,q\begin{aligned} \eta_{s,i}={}&\mu_s+\alpha_r+\alpha_j +(d_r+\gamma_r^\top x_i)\tau_i+(d_j-d_r)D(\tau_i)\\ &+b_j^\top x_i+\beta_L h_i\\ &+d_i\left(\beta_{\rm def}+f^\top x_i+d_{\rm def}\tau_i\right)+u_{s,q} \end{aligned}
αr=σα,regzr,αj=σα,ctryzjbr=b0+σb,regzb,r,bj=br(j)+σb,ctryzb,jdr=d0+σd,regzd,r,dj=dr(j)+σd,ctryzd,jγr=γ0+σγzγ,r,us,q=σu,szu,s,q\begin{gathered} \alpha_r=\sigma_{\alpha,\rm reg}z_r,\qquad \alpha_j=\sigma_{\alpha,\rm ctry}z_j\\ b_r=b_0+\sigma_{b,\rm reg}z_{b,r},\qquad b_j=b_{r(j)}+\sigma_{b,\rm ctry}z_{b,j}\\ d_r=d_0+\sigma_{d,\rm reg}z_{d,r},\qquad d_j=d_{r(j)}+\sigma_{d,\rm ctry}z_{d,j}\\ \gamma_r=\gamma_0+\sigma_\gamma z_{\gamma,r},\qquad u_{s,q}=\sigma_{u,s}z_{u,s,q} \end{gathered}

Smoking has regional and country age profiles and trends. The regional term γ allows the trend to vary with age. The current-use adjustment includes an intercept, age terms f and a time term.

Cigarette share among smokers

ηc,i=μc+αc,j+dc,0τi+(dc,jdc,0)D(τi)+βc,defdi+bc,0xi+uc,q\eta_{c,i}=\mu_c+\alpha_{c,j}+d_{c,0}\tau_i +(d_{c,j}-d_{c,0})D(\tau_i)+\beta_{c,\rm def}d_i+b_{c,0}^{\top}x_i+u_{c,q}
αc,j=σc,regzc,r(j)+σc,ctryzc,jdc,j=dc,0+σd,czd,c,j,uc,q=σu,czu,c,q\begin{gathered} \alpha_{c,j}=\sigma_{c,\rm reg}z_{c,r(j)}+\sigma_{c,\rm ctry}z_{c,j}\\ d_{c,j}=d_{c,0}+\sigma_{d,c}z_{d,c,j},\qquad u_{c,q}=\sigma_{u,c}z_{u,c,q} \end{gathered}

Cigarette share has regional and country intercepts, a global age profile and country departures from a global trend. It has no regional trend parameter.

Smokeless use among non-smokers

η,i=μ+α,j+d,0τi+(d,jd,0)D(τi)+β,defdi+b,rxi+β,Lhi+u,q\begin{aligned} \eta_{\ell,i}={}&\mu_\ell+\alpha_{\ell,j}+d_{\ell,0}\tau_i +(d_{\ell,j}-d_{\ell,0})D(\tau_i)\\ &+\beta_{\ell,\rm def}d_i+b_{\ell,r}^{\top}x_i+\beta_{\ell,L}h_i+u_{\ell,q} \end{aligned}
α,j=σ,regz,r(j)+σ,ctryz,jd,j=d,0+σd,zd,,j,u,q=σu,zu,,qb,r={b,0+σ,bz,b,r,rE,b,0,rE.\begin{gathered} \alpha_{\ell,j}=\sigma_{\ell,\rm reg}z_{\ell,r(j)}+\sigma_{\ell,\rm ctry}z_{\ell,j}\\ d_{\ell,j}=d_{\ell,0}+\sigma_{d,\ell}z_{d,\ell,j},\qquad u_{\ell,q}=\sigma_{u,\ell}z_{u,\ell,q}\\ b_{\ell,r}=\begin{cases} b_{\ell,0}+\sigma_{\ell,b}z_{\ell,b,r},&r\in E,\\ b_{\ell,0},&r\notin E. \end{cases} \end{gathered}

The set E contains Southern Asia, Northern Europe, Southeastern Asia, North Africa and Middle East, Sub-Saharan Africa, and Oceania and Pacific in the model-region mapping. Only these six regions have regional departures from the global smokeless age profile. Other regions use the global profile. The definition adjustment is estimated with a prior centered on 0.3 times the smoking definition adjustment.

Prior distributions

Normal distributions use mean and standard deviation. Exponential and gamma distributions use rate parameters. Lognormal distributions use mean and standard deviation on the log scale. Vector priors apply independently to each component. Prior factors are independent except for the conditional smokeless definition adjustment shown below. All z coefficients have N(0, 1) priors.

ParameterPrior
μs\mu_sN(2,1)N(-2,1)
βL\beta_LN(0.02,0.05)N(-0.02,0.05)
βdef\beta_{\rm def}N(0.3,1)N(0.3,1)
σ\sigmaLognormal(log0.7,0.5)\operatorname{Lognormal}(\log 0.7,0.5)
ν\nuGamma(10,1),ν4\operatorname{Gamma}(10,1),\quad\nu\ge4
σα,reg\sigma_{\alpha,\rm reg}Exp(2)\operatorname{Exp}(2)
σα,ctry\sigma_{\alpha,\rm ctry}Exp(3)\operatorname{Exp}(3)
σu,s\sigma_{u,s}Exp(3)\operatorname{Exp}(3)
b0b_0N(0,1.5)N(0,1.5)
σb,reg\sigma_{b,\rm reg}Exp(2.5)\operatorname{Exp}(2.5)
σb,ctry\sigma_{b,\rm ctry}Exp(3)\operatorname{Exp}(3)
d0d_0N(0.02,0.05)N(-0.02,0.05)
σd,reg\sigma_{d,\rm reg}Exp(3)\operatorname{Exp}(3)
σd,ctry\sigma_{d,\rm ctry}Exp(3)\operatorname{Exp}(3)
γ0\gamma_0N(0,0.05)N(0,0.05)
σγ\sigma_\gammaExp(5)\operatorname{Exp}(5)
ffN(0,0.15)N(0,0.15)
ddefd_{\rm def}N(0,0.03)N(0,0.03)
μc\mu_cN(1.5,1)N(1.5,1)
σc,reg\sigma_{c,\rm reg}Exp(2)\operatorname{Exp}(2)
σc,ctry\sigma_{c,\rm ctry}Exp(3)\operatorname{Exp}(3)
bc,0b_{c,0}N(0,0.75)N(0,0.75)
dc,0d_{c,0}N(0,0.03)N(0,0.03)
σd,c\sigma_{d,c}Exp(5)\operatorname{Exp}(5)
βc,def\beta_{c,\rm def}N(0,0.5)N(0,0.5)
σu,c\sigma_{u,c}Exp(4)\operatorname{Exp}(4)
μ\mu_\ellN(4,1)N(-4,1)
σ,reg\sigma_{\ell,\rm reg}Exp(2)\operatorname{Exp}(2)
σ,ctry\sigma_{\ell,\rm ctry}Exp(3)\operatorname{Exp}(3)
b,0b_{\ell,0}N(0,0.75)N(0,0.75)
σ,b\sigma_{\ell,b}Exp(3)\operatorname{Exp}(3)
β,L\beta_{\ell,L}N(0.02,0.015)N(-0.02,0.015)
d,0d_{\ell,0}N(0,0.05)N(0,0.05)
σd,\sigma_{d,\ell}Exp(5)\operatorname{Exp}(5)
β,defβdef\beta_{\ell,\rm def}\mid\beta_{\rm def}N(0.3βdef,0.2)N(0.3\beta_{\rm def},0.2)
σu,\sigma_{u,\ell}Exp(4)\operatorname{Exp}(4)
ρ\rhoN(0.5,0.25),ρ>0N(0.5,0.25),\quad\rho>0
τhet\tau_{\rm het}Lognormal(log0.2,0.5)\operatorname{Lognormal}(\log 0.2,0.5)

Survey uncertainty

The model uses a Student-t distribution on logit-transformed prevalence. Its heavier tails reduce the influence of observations far from the fitted curve. Survey effects account for differences shared by observations from the same survey instance.

Reported standard errors are used when available. Otherwise, positive sample sizes provide a binomial approximation. Where neither is available, the observation scale depends on a weight based on age-range width and sample size. Most standard errors used in this fit are approximations, which may not capture a complex survey design.

Likelihood and observation weights

Reported percentages vᵢ are converted to proportions pᵢ. Values below 0.001 or above 0.999 are set to those bounds before taking the logit, including reported 0% and 100%. This adjustment can matter where many surveys report very low prevalence.

pi=vi100,yi=logit{min[0.999,max(0.001,pi)]}p_i=\frac{v_i}{100},\qquad y_i=\operatorname{logit}\{\min[0.999,\max(0.001,p_i)]\}
mi={ηs,i,ki=s,logit(sici),ki=c,logit{si+(1si)i},ki=a,yiθ,u  tν(mi,ai),ai={SE~logit,i2+τhet2,SE,σ/w~i,no SE,SD(yiθ,u)=aiνν2,L(θ,u)=iftν(yimi,ai)\begin{gathered} m_i=\begin{cases}\eta_{s,i},&k_i=s,\\ \operatorname{logit}(s_i c_i),&k_i=c,\\ \operatorname{logit}\{s_i+(1-s_i)\ell_i\},&k_i=a, \end{cases}\\ y_i\mid\theta,u\ \sim\ t_\nu(m_i,a_i),\qquad a_i=\begin{cases}\sqrt{\widetilde{\rm SE}_{\rm logit,i}^{\,2}+\tau_{\rm het}^2},&\text{SE},\\ \sigma/\sqrt{\widetilde w_i},&\text{no SE}, \end{cases}\\ \operatorname{SD}(y_i\mid\theta,u)=a_i\sqrt{\frac{\nu}{\nu-2}},\qquad \mathcal L(\theta,u)=\prod_i f_{t_\nu}(y_i\mid m_i,a_i) \end{gathered}

The product kᵢ selects the likelihood location mᵢ. The scale aᵢ is not the Student-t standard deviation. With a usable standard error, aᵢ combines survey uncertainty and residual heterogeneity τhet. Without one, aᵢ uses the normalized observation weight and has no additional heterogeneity floor. Composed cigarette and any-tobacco probabilities are bounded away from zero and one by 10⁻¹² for numerical stability.

qi=max{pi(1pi),0.001(10.001)}SEp,i={ei/100,absolute SE ei>0,eipi/100,relative SE ei>0,qi/ni,no reported SE, ni>0,NA,otherwise,SE~logit,i=min ⁣(SEp,iqi,2)\begin{gathered} q_i^*=\max\{p_i(1-p_i),0.001(1-0.001)\}\\ {\rm SE}_{p,i}=\begin{cases} e_i/100,&\text{absolute SE }e_i>0,\\ e_i p_i/100,&\text{relative SE }e_i>0,\\ \sqrt{q_i^*/n_i},&\text{no reported SE, }n_i>0,\\ \text{NA},&\text{otherwise}, \end{cases}\\ \widetilde{\rm SE}_{\rm logit,i}=\min\!\left(\frac{{\rm SE}_{p,i}}{q_i^*},2\right) \end{gathered}
wi=1ri+1{min(ni,10000)/1000,ni>0,1,otherwise,w~i=wimean(w)w_i^*=\frac{1}{r_i+1}\begin{cases} \sqrt{\min(n_i,10000)/1000},&n_i>0,\\ 1,&\text{otherwise}, \end{cases}\qquad \widetilde w_i=\frac{w_i^*}{\operatorname{mean}(w^*)}

Here eᵢ is the reported standard error, nᵢ the sample size and rᵢ the corrected age-range width. Absolute standard errors are in percentage points. Four survey series report relative standard errors instead, requiring multiplication by prevalence. These are Australia in 2012, Spain in 2020, the Philippines in 2003 and Thailand in 2021. Logit standard errors are capped at two. Fallback weights have mean one within each sex.

Residual errors are independent conditional on model parameters and survey effects. The model has no sampling-covariance matrix for overlapping age ranges or nested tobacco indicators.

Population summaries

Crude prevalence
The estimated percentage of adults who use the selected product. Ages are weighted by the country’s actual adult population.
Age-standardized prevalence
A weighted average using the same reference age distribution for every country and year. It supports comparisons that are less affected by differences in population age structure.
Number of users
The sum of age-specific prevalence multiplied by the corresponding population. Counts are expressed in people.
Point-age estimates
Values at ages 15, 20, 25 and so on through 100. These refer to exact ages, not five-year age-group averages.

Population weights come from UN World Population Prospects 2022. Both-sex crude prevalence weights females and males by their adult populations. Both-sex age-standardized prevalence gives each sex equal weight. Regional and global results weight countries by their adult populations.

Hong Kong, Macao, Taiwan and Kosovo lack population weights in this input. Their crude prevalence, user counts and endgame probabilities are unavailable. Their sex-specific age-standardized estimates and point-age curves remain available. Population-weighted regional and global summaries cover the other 193 countries and territories.

Aggregation equations

For posterior draw k, Nⱼₐₜ is the population in people and pⱼₐₜ is prevalence at age a in country j and year t, for one sex and indicator. P denotes crude prevalence, C the number of users and Pstd age-standardized prevalence. The age-100 evaluation represents the population aged 100 and older. Survey effects are set to zero when generating these trajectories.

Pjt(k)=a=15100Njatpjat(k)a=15100Njat,Cjt(k)=a=15100Njatpjat(k)Pstd,jt(k)=a=15100vapjat(k),a=15100va=1\begin{gathered} P_{jt}^{(k)}=\frac{\sum_{a=15}^{100}N_{jat}p_{jat}^{(k)}}{\sum_{a=15}^{100}N_{jat}},\qquad C_{jt}^{(k)}=\sum_{a=15}^{100}N_{jat}p_{jat}^{(k)}\\ P_{{\rm std},jt}^{(k)}=\sum_{a=15}^{100}v_a p_{jat}^{(k)},\qquad \sum_{a=15}^{100}v_a=1 \end{gathered}

The reference age weights vₐ are normalized over adult ages. Each five-year band from 15–19 to 80–84 is split equally across its five integer ages. The reference weight for ages 85 and older is split equally across ages 85 to 100. This final allocation is a convention of the analysis.

Pboth,jt(k)=Nm,jtPm,jt(k)+Nf,jtPf,jt(k)Nm,jt+Nf,jtPstd,both,jt(k)=12(Pstd,m,jt(k)+Pstd,f,jt(k)),Cboth,jt(k)=Cm,jt(k)+Cf,jt(k)Δjt(k)=Pm,jt(k)Pf,jt(k)\begin{gathered} P_{{\rm both},jt}^{(k)}=\frac{N_{m,jt}P_{m,jt}^{(k)}+N_{f,jt}P_{f,jt}^{(k)}}{N_{m,jt}+N_{f,jt}}\\ P_{{\rm std,both},jt}^{(k)}=\tfrac12\left(P_{{\rm std},m,jt}^{(k)}+P_{{\rm std},f,jt}^{(k)}\right),\qquad C_{{\rm both},jt}^{(k)}=C_{m,jt}^{(k)}+C_{f,jt}^{(k)}\\ \Delta_{jt}^{(k)}=P_{m,jt}^{(k)}-P_{f,jt}^{(k)} \end{gathered}
PG,t(k)=jGNjtPjt(k)jGNjt,Pstd,G,t(k)=jGNjtPstd,jt(k)jGNjtP_{G,t}^{(k)}=\frac{\sum_{j\in G}N_{jt}P_{jt}^{(k)}}{\sum_{j\in G}N_{jt}},\qquad P_{{\rm std},G,t}^{(k)}=\frac{\sum_{j\in G}N_{jt}P_{{\rm std},jt}^{(k)}}{\sum_{j\in G}N_{jt}}

Male and female draws are paired as independent fitted distributions. Their counts add, and their prevalence difference Δ is calculated, within each pair. For a group of countries G, standardized country draws are weighted by adult population totals before the two sex-specific aggregates receive equal weight. Aggregation precedes every median, interval or target calculation. Combining published medians or interval endpoints would not give the same result.

Endgame probabilities

The endgame threshold is crude tobacco-use prevalence below 5% among adults aged 15 and older. It is evaluated for the selected country, population, product and use definition. A prevalence of exactly 5% does not meet this threshold.

Below 5% in the selected year
The proportion of model draws below the threshold in that year.
Below 5% at least once by that year
The proportion that reach the threshold at least once from 2000 through that year. Prevalence can subsequently rise above 5%, so this probability can be higher.

For example, an in-year probability of 80% means that 800 of the 1,000 model draws fall below 5% in that year. It is a probability of meeting the threshold, not the percentage of people who use tobacco. The both-sex probability is calculated from combined-sex prevalence in each draw, rather than averaging the female and male probabilities.

The separate relative-reduction benchmark is at least a 30% reduction from 2010 prevalence. It is evaluated against each draw’s own 2010 value. Compare probabilities only when the product, population, year and benchmark match.

Threshold and first-crossing equations

H is one when a draw meets the chosen threshold and zero otherwise. K = 1,000 is the number of draws. The average π̂ₜ is the in-year probability.

Hrel,t(k)=1{t2010}1{Pt(k)0.7P2010(k)}Habs,t(k)=1{Pt(k)<0.05},π^t=1Kk=1KHt(k)\begin{gathered} H_{{\rm rel},t}^{(k)}=\mathbf 1\{t\ge2010\}\,\mathbf 1\{P_t^{(k)}\le0.7P_{2010}^{(k)}\}\\ H_{{\rm abs},t}^{(k)}=\mathbf 1\{P_t^{(k)}<0.05\},\qquad \widehat\pi_t=\frac1K\sum_{k=1}^K H_t^{(k)} \end{gathered}
T={2000,2001,,2040},Ak=min{uT:Hu(k)=1},min=F^(t)=1Kk=1K1{Akt},Pr^(A>2040)=1F^(2040)Qα=min{tT:F^(t)α},α{0.025,0.5,0.975}\begin{gathered} \mathcal T=\{2000,2001,\ldots,2040\},\qquad A_k=\min\{u\in\mathcal T:H_u^{(k)}=1\},\quad \min\varnothing=\infty\\ \widehat F(t)=\frac1K\sum_{k=1}^K\mathbf 1\{A_k\le t\},\qquad \widehat{\Pr}(A>2040)=1-\widehat F(2040)\\ Q_\alpha=\min\{t\in\mathcal T:\widehat F(t)\ge\alpha\},\qquad \alpha\in\{0.025,0.5,0.975\} \end{gathered}

Aₖ is the first qualifying year for draw k on the annual grid from 2000 to 2040. The achieved-by probability F̂(t) includes all draws that crossed by t, including those that later rebound. Relative reduction cannot be achieved before 2010. A draw already below 5% in 2000 is recorded as crossing in 2000, without inferring an earlier date.

Draws that do not cross by 2040 remain in the denominator. An achievement-year median is unavailable if fewer than half the draws cross by then. The same rule applies to other quantiles. No crossing by 2040 does not mean no crossing at any future time.

Fitting and uncertainty intervals

The model was fitted with Stan’s No-U-Turn sampler. Each sex used four chains, with 1,000 warmup and 1,000 retained iterations per chain. Publication calculations use 1,000 sampled draws per sex. The displayed estimate is the posterior median. The 95% credible interval runs from the 2.5th to the 97.5th percentile.

Intervals describe uncertainty in the fitted prevalence trajectory under the model assumptions. They are pointwise intervals, so they apply separately to each reported age or year. They do not include the extra variability of a new survey, uncertainty in population weights, or future shocks.

Sampling diagnostics

R-hat measures agreement across chains. Values closer to one are preferable. Effective sample size measures how much independent sampling information the correlated draws provide. The table reports the worst value of each measure across parameters.

PopulationHighest R-hatLowest bulk ESSDivergences
Females1.031880
Males1.122250

No divergent transitions were recorded. Some parameters nevertheless mixed poorly across chains, particularly in the male fit. The diagnostics do not show uniformly reliable sampling. Completed checks include posterior predictive checks and approximate leave-one-out assessment of the smoking component. A completed assessment that withholds future years is not available for these fits.

Limits of the estimates

  • Data are uneven across countries, ages and years. Estimates without local observations rely on the hierarchy and priors. A smooth curve is not evidence of complete survey coverage.
  • Survey uncertainty is approximated, overlapping observations are not fully independent, and very low reported prevalence is adjusted before fitting. These choices can affect the results.
  • Females and males are fitted independently. Combined results do not account for shared sampling errors or future shocks across sexes. Birth-cohort labels are calculated as year minus age. The model does not estimate a separate cohort effect.
  • Some parameters have weak sampling diagnostics. Forecast damping is determined by its prior, and forecast performance has not been established by withholding future years. Endgame probabilities depend on these assumptions and are not guarantees or estimates of a policy’s causal effect.