What's in and out. This version uses only already-public material: two commissioned public evaluations, the public workshop record, and the public post-workshop survey. Our structured belief elicitation is still under respondent review, so its numbers are deliberately excluded here (see section 8).
1. What this decision is about
A funder comparing a psychotherapy or mental-health programme against a bednet, a vaccine, or a cash transfer has to put both on one scale. In practice that often means converting between DALYs[1]Disability-adjusted life year: a year of healthy life lost to death or disability, weighted by condition-specific disability weights. The standard unit of global-health cost-effectiveness. averted and WELLBYs[2]Wellbeing-adjusted life year: one point on a 0-to-10 life-satisfaction scale, for one person, for one year. gained.
This is the operational form of one of our Wellbeing Pivotal Questions. The canonical formulation we put to participants (WELL_01) asks what combination of subjective-wellbeing survey data, income and health-outcome data, derived metrics, and conversions between them would be best for choosing between interventions affecting mental health, physical health, or consumption.[3]Full text and the related DALY_01-05 formulations are on our public Coda page; the elicitation form states each question as it was put to participants.
Scope/audience: The two conversion pipelines below are used by effective-altruism-adjacent global health and development funders doing explicit cross-domain cost-effectiveness analysis. They are not standard practice in health economics generally, where cost-per-QALY within a domain is the norm and cross-domain wellbeing conversion is rare.
Within that group, the two pipelines a funder is most likely to meet — Founders Pledge's moral weights and the Happier Lives Institute's cost-effectiveness work, which the Founders Pledge weights build on — both pass through the same step: treating one standard deviation of improvement on a depression or anxiety instrument[4]A standard questionnaire such as the PHQ-9 for depression or the GAD-7 for anxiety, scored and compared in standard-deviation units. as one standard deviation of improvement in life satisfaction. Anyone building on either analysis inherits this step, whether they know it or not.
So the practical question, and the one this page tries to answer, is narrow: can I keep using this step, and how much could it distort my rankings?
Where this page comes from, and what it's for
The Unjournal commissions public expert evaluations of impactful research; our Pivotal Questions workshops work backwards from decisions funders and policy actors are actually facing. Founders Pledge raised the WELLBY-reliability and DALY↔WELLBY questions with us through that programme.
As an initial response we commissioned two public evaluations of Benjamin, Cooper, Heffetz, Kimball & Zhou's scale-use paper, wrote two framing analyses (linear WELLBY reliability; DALY↔WELLBY conversion), convened a March 2026 workshop with the paper's authors, HLI, the evaluators, and practitioners from Founders Pledge and Coefficient Giving, and ran a structured belief elicitation that is now under respondent review. This page distills the practical upshot for a funder; section 3 lists the sources and how we attribute claims to them.
2. Bottom line
The five points below are our current reading, not findings. They are our synthesis of the workshop, the commissioned evaluations, and follow-up correspondence. They have not been reviewed or endorsed by the workshop participants, by the evaluators, or by our research contractors — we plan to put this draft to them shortly, and expect some of it to change. Where a point rests on one person's argument rather than on evidence, we say so in the section that argues it.
Reasoning and sources for each point are in sections 4–7, linked from each item.
- The SD-to-SD step is defensible as a working assumption, but it is nobody's preferred method — including its originators' (§4). Keep using it, state it as an assumption, and test how much your ranking depends on it, rather than presenting the SD equivalence itself as established.
- The support for it is thinner than the "near-1:1 exchange rate" in current use might suggest. That rate comes out of HLI's SD-based conversion work, described by HLI and Founders Pledge at the workshop. It isn't purely assumed — the conversion work rests on data linking mental-health and life-satisfaction measures, and there is some regression evidence on how the two move together — but that evidence is limited, and we still need to pin down exactly what it shows. [To check in review: the exact empirical basis and figure behind the near-1:1 rate.] However, conversion is a ratio operation, and the evaluations we commissioned suggest ratios are where scale-use problems bite hardest (§5).
- The error is plausibly directional rather than symmetric. To be precise about the comparison: the concern is that converting a mental-health-instrument effect into life-satisfaction units at 1 SD = 1 SD will tend to produce a larger wellbeing estimate than directly measuring the same intervention's effect on life satisfaction would — and direct measurement is the benchmark everyone at the workshop said they would prefer (§4). The arguments we have seen mostly push in that direction, rather than scattering symmetrically.
The intuition, and how confident we are in this
A depression instrument tracks the symptoms a therapy targets. A life-satisfaction question averages over everything in a person's life — health, income, relationships — most of which the therapy did not touch. An improvement concentrated in the targeted symptoms will generally register more strongly on the targeted instrument than in overall life satisfaction, so equating the two SDs will tend to favor mental-health interventions.
Compression and ceiling arguments pointing the same way have been made in funders' published assessments of psychotherapy charities — we believe GiveWell's StrongMinds assessment is the main public source. [To check in review: the exact form of the argument there, and whether any direct empirical evidence backs the direction — as stated here it is an argument, not a measurement.]
Epistemic status: this is our reading of arguments raised at the workshop and in those published assessments. Plausible, not established, and not a workshop consensus. It is also the point we would retract fastest if the crosswalk data (§7) came out the other way.
- DALY disability weights and observed life-satisfaction losses do not line up across health conditions. In the comparisons Samuel Dupret presented, mental-health conditions show life-satisfaction losses that are large relative to their disability weights, while many physical conditions show the reverse. If that holds, a single global conversion factor over-weights some conditions and under-weights others by construction (§5), and domain-specific factors may be more promising than a more precise global one.
- Some fixes are small additions to existing instruments. Two vignette calibration questions[5]Short descriptions of hypothetical people that respondents rate on the same scale. Differences in how people rate identical vignettes reveal differences in how they use the scale, which is what allows a correction. is the minimum implementation of the scale-use correction in the paper we had evaluated (§6). We don't have actual cost estimates — the claim is only that two added survey items are small relative to the cost of fielding a trial. And none of this has been validated in intervention settings, which is itself a key research gap (§7).
3. What this is based on
| Component | Status |
|---|---|
| Commissioned public evaluations of Benjamin, Cooper, Heffetz, Kimball & Zhou, "Adjusting for Scale-Use Heterogeneity in Self-Reported Well-Being" (NBER WP 31728) — evaluators Caspar Kaiser and Alberto Prati | Public, with DOI: evaluation package |
| Framing analyses: linear WELLBY reliability; DALY↔WELLBY conversion | Public, on this site |
| Workshop, 16 March 2026 — paper authors (Benjamin, Heffetz, Kimball), Happier Lives Institute (Plant, McGuire, Dupret), Kaiser, Julian Jamison, Dean Jamison, and practitioners from Founders Pledge (Lerner) and Coefficient Giving (Hickman) | Public edited transcript and video; full transcript |
| Post-workshop participant survey | Public: survey results |
| Background reading assembled for the workshop | Public: readings |
| Structured belief elicitation, 8 respondents (linear-WELLBY reliability, WELLBYs per DALY, adoption forecasts) | Under respondent review — not reported here |
On attribution: where a claim rests on what one person said, we name them and link the spot in the transcript; where it comes from the commissioned evaluations, we link those; and where it is our own synthesis we say so ("our reading"). Weight them accordingly. The transcript is lightly edited from the Zoom recording and automated captions, so check quotes against the full transcript or the recording before relying on exact wording.
4. Nobody's preferred method — including its originators'
Founders Pledge arrived at its WELLBY/DALY conversion by triangulation, as Matt Lerner described at the workshop: WELLBY to income-doubling (from Joel McGuire's work), income-doubling to lives-saved, lives-saved to DALYs, then backing out the remaining side. In his words: "a bit kludgy — we had two sides of a triangle and filled in the third for interconvertibility."
Where the SD-to-SD step sits in that triangle
Spelling the triangle out: leg one converts WELLBYs to income-doublings, using McGuire's conversion work; leg two converts income-doublings to lives saved; leg three converts lives saved to DALYs. The DALY↔WELLBY rate is then whatever number makes the triangle close.
The SD-to-SD step sits inside leg one. McGuire's work pools trials that measured different things — depression scores, psychological distress, life satisfaction — by putting each on a standard-deviation scale and treating those SDs as interchangeable. So if one SD on a depression scale actually corresponds to, say, only 0.7 SD of life satisfaction, leg one is off by roughly that factor, and the DALY↔WELLBY rate backed out at the end is off with it. That is why a funder can accept the triangulation logic and still want to know how much weight the SD leg can bear.
The step exists because of data constraints. Samuel Dupret (HLI) explained that HLI lacks rich 0–10 scale datasets for LMIC interventions, so it converts trial results to standard deviations, then to a 0–10 scale using typical Cantril Ladder[6] SDs from the World Happiness Report — "This isn't because we love SD conversions — it's data constraints." Michael Plant: "We're being data omnivores." And Lerner, agreeing: "If we had life satisfaction data for everything, we'd skip the SD conversion."
The consensus itself is worth knowing: the people who built this conversion, and the people who use it, all describe it as a stopgap. Treat the number that way.
5. Three problems that a better estimate of the exchange rate won't fix
The SD units import properties of whichever samples were involved
An effect measured "in standard deviations" is the raw change divided by how spread out the outcome was in that particular sample. The same raw improvement therefore looks larger, in SD terms, when the trial sample was more uniform, and smaller when it was more varied. The pipeline then converts back to a 0–10 scale by multiplying by the spread of a different population altogether — in HLI's implementation, country-level Cantril SDs from the World Happiness Report, per Dupret's description.
So a comparison between two interventions can shift with the variance of the samples each happened to be studied in, at both ends of the conversion. How much of a problem this is, is genuinely open. There is a counterargument: normalising by a sample's spread is not obviously wrong, since a context with little variation in an outcome may also be one where a given raw change means more — so SD scaling could be capturing something real about the context rather than just an artefact. How far variance differences reflect that, versus sampling and measurement choices with nothing to do with the intervention, is a methodological question we would like dug into properly in review. At minimum, the choice is unmodelled and invisible in the headline number. (Our reading — a version of a familiar concern about standardised effect sizes in meta-analysis; the WHR-SD step itself is HLI's own description of their method.)
Depression instruments and life-satisfaction measures overlap without nesting
Depression and anxiety instruments track negative affect and day-to-day functioning; a life-satisfaction question asks for a global evaluation of your life. The two are related, but neither contains the other. Why that matters here: even if one SD tracked one SD well in general population surveys, a treatment effect on one need not map one-for-one onto the other. A therapy that relieves the symptoms an instrument asks about may move a person's overall life evaluation much less, because the rest of their circumstances have not changed. Treating the two SDs as interchangeable assumes this away.
The implication, spelled out: measuring an intervention with the instrument that targets its symptoms will tend to overstate its effect, even in SD terms, relative to measuring the same intervention's effect on life satisfaction. So SD-to-SD comparison will tend to favor interventions evaluated on narrowly targeted instruments (a therapy on a depression scale) over interventions evaluated on broad measures (a cash transfer on life satisfaction). It plausibly cuts the other way too: a broad intervention scored on a narrow instrument would have much of its effect missed.
Where this comes from: that these instruments measure related but distinct constructs is uncontroversial in the measurement literature — the point is not original to us. The implication for cross-intervention conversion drawn here is our own synthesis, with related points in the Kimball discussion (neither single measure captures everything people care about). It is not a workshop conclusion, and it is one of the claims we would most like challenged in review.
Ratios are far more sensitive than the coefficients underneath them
Caspar Kaiser's evaluation of the Benjamin et al. results makes a point that bears directly on funders: correcting for scale use moves individual coefficients only modestly, but their ratios can move a lot, because modest shifts in numerator and denominator compound. Cross-intervention comparison is a ratio operation, so funders sit exactly where the method is most fragile. Miles Kimball's related example: scale-use correction moves the income coefficient by roughly a factor of five in their unemployment application, which is first-order for anything routed through income-doubling equivalences, including the Founders Pledge triangle above.
How much this matters in intervention settings specifically is close to unknown. Kaiser's assessment is that we have essentially no evidence on scale use in trials, rich or poor country; the literature is population surveys.
6. What we recommend
Do now, at no cost
- State the SD conversion as an assumption wherever a comparison depends on it, and show bounds rather than a single number.
A minimal way to do the bounds
Pick a low and a high value for the contested step — the published DALY↔WELLBY range runs roughly 2–15; see the anchors and interactive calculator on our conversion page — recompute your ranking at both ends, and report whether the decision flips. If it does not flip, say so and move on. If it does, the conversion factor is doing real work in your decision and deserves real attention.
- Where a trial measured life satisfaction directly, use that in preference to a converted mental-health measure. This is the one recommendation everyone at the workshop endorsed; it is what HLI and Founders Pledge each said they would do if the data allowed (§4).
- Flag when a comparison hinges on the neutral point[7]The life-satisfaction level at which a year of life counts as neither adding to nor subtracting from wellbeing. Life-saving versus life-improving comparisons can hinge on where you put it. — see the discussion on our linear WELLBY page.
- Flag when a comparison hinges on treating the 0–10 scale as linear, so that a move from 3 to 4 counts the same as 7 to 8. Kimball's view at the workshop was that scale curvature is partly an ethical choice about inequality aversion, not only a measurement fact; our linear WELLBY analysis covers this.[8]
- Do not compare a mental-health intervention against a physical-health one using a single global factor without saying that is what you did (bottom line, point 4).
Do if you have budget or control an instrument
- Add two vignette calibration questions to any wellbeing instrument you fund — the minimum implementation of the Benjamin et al. correction, for which the authors have published a practical guide.[9] Visual calibrations give partial correction where vignettes do not fit. Small in survey-cost terms, though fielding and validating them properly takes real work.
- For any psychological-intervention trial, collect calibration before and after treatment, and in controls. Otherwise a measured gain could be a genuine welfare change or a therapy-induced change in how people use the scale, and you cannot tell which. This gap is §7's top item.
- Add stated-preference tradeoff questions so the exchange rate can be estimated rather than assumed. Dan Benjamin's suggestion in follow-up discussion: absent extra data, comparing SD improvements is hard to beat, but a small number of added questions would let you estimate the rate directly.
Don't bother yet
- Full multi-dimensional wellbeing indices in LMIC field settings, absent dedicated funding.
- Chasing a more precise single global DALY↔WELLBY factor. If the divergence in point 4 is real, the global factor is the wrong object to refine.
7. What would change which part of this
Ranked by our judgment of value of information; each item names the bottom-line point it bears on.
- Calibration plus response-shift diagnostics inside a mental-health intervention trial, separating welfare change from therapy-induced change in scale use. Bears on point 1 and point 3: it directly tests whether measured mental-health gains are the kind of thing the SD step assumes. Raised independently by Kaiser and Kimball; as far as we can tell, a genuine literature gap.
- Calibration and value elicitation in a globally low-income population — the biggest unknown behind point 2, since scale-use evidence comes almost entirely from rich-country population surveys.
- A crosswalk harvest: systematically collect trials that report both a depression or anxiety instrument and life satisfaction, and extract the distribution of the implied ratio. This would replace the assumed exchange rate with a measured one, answering the question this brief is about — how many WELLBYs per DALY, for mental health — better than any single number, and extending rather than duplicating HLI's conversion work.
Why a distribution rather than a better point estimate
If the ratio varies a lot across conditions, populations, and instruments — which points 2 and 4 suggest — then the useful output is the spread itself. It tells a funder when the conversion is safe, when it is doing heavy lifting, and what sensitivity bounds to use. A single pooled number would hide exactly the variation that matters.
- Direct validation of the linear WELLBY against better-specified alternatives, which bears on the cardinality concern in §6. Concretely: run the single 0–10 question alongside calibrated items, multi-item batteries, and tradeoff questions in the same sample, and check whether they rank the same interventions the same way. Per Benjamin, this has not been done.
We are looking for one funder or research group to take on item 3. It is comparatively small, it uses already-published data, and it converts an assumption into a measurement. If that is you, write to us.
8. What is still coming
Elicitation results, in a later version. Eight participants gave central estimates with 80% credible intervals on linear-WELLBY reliability, WELLBYs per DALY, and adoption forecasts. Respondents are reviewing a corrected analysis and we will publish once that round closes. Preview, without numbers: broad support for WELLBY-style data as useful, considerably less for a single linear life-satisfaction score as the endpoint, and genuine unresolved spread on the conversion factor. The spread is arguably the finding, and reporting a tidy central value would misrepresent where things stand.
A limitation in our elicitation, and why we can't fully repair it
The sliders in our elicitation form displayed a visible starting value that could be submitted without being touched. The form recorded only the final value, not whether the respondent actually moved the slider. So an answer sitting exactly at the default could be a deliberate answer that happened to match, an anchored answer, or a slider that was never touched, and we cannot distinguish these after the fact.
What we can and will do: report results with and without exact-default answers as a sensitivity check, and say how many responses fall in that category. What we cannot do is know which of them were genuine. Our later workshop forms start from blank sliders, which removes the problem at the source. We would rather publish this limitation than quietly drop it.
9. How to use and cite this
Written for funders and CEA practitioners doing cross-domain comparisons of the kind described in section 1, who must choose between interventions measured in different units now, without waiting for the literature to settle. It is not a verdict on whether wellbeing measures are valid, and it is not addressed to health economists working within a single domain.
Please don't circulate or cite it while the working-draft banner is up. Comments and disagreement are very welcome: select any text to annotate via Hypothes.is, including the notes below, or email contact@unjournal.org.
Notes and sources
- Disability-adjusted life year: a year of healthy life lost to death or disability, weighted by condition-specific disability weights, and the standard unit of global-health cost-effectiveness. On its construction and known criticisms see Cookson et al., Quality-adjusted life years (workshop reading). ↩
- Wellbeing-adjusted life year: one point on a 0–10 life-satisfaction scale, for one person, for one year. For the policy framework that popularised it see HM Treasury's Green Book wellbeing guidance and Layard et al., When to release the lockdown (both workshop readings). ↩
- The canonical formulations (WELL_01–07, DALY_01–05) are on our public Wellbeing PQ page on Coda. The elicitation form states each question as it was put to participants, with the canonical text shown alongside. ↩
- Typically the PHQ-9 (depression) or GAD-7 (anxiety), scored and compared in standard-deviation units. The mapping between such instruments and life-satisfaction scales is precisely what is at issue in §5. ↩
- Vignettes are short descriptions of hypothetical people that respondents rate on the same scale they rate themselves on. Because every respondent rates the same vignettes, differences in those ratings identify differences in scale use, which is what permits a correction. Method: Benjamin, Cooper, Heffetz, Kimball & Zhou, Adjusting for scale-use heterogeneity; practical implementation: the authors' guide to calculating benchmarks using calibration questions. Our commissioned evaluations assess how far the method carries. ↩
- The Cantril Ladder is the 0–10 "ladder" life-evaluation question used in the Gallup World Poll and reported in the World Happiness Report. On whether such ordinal scales support the cardinal comparisons made of them, see Bond & Lang, The sad truth about happiness scales, and Kaiser & Vendrik, How threatening are transformations of happiness scales? (both workshop readings, and they reach different conclusions). ↩
- The neutral point is the life-satisfaction level at which a year of life counts as neither adding to nor subtracting from wellbeing. It has no effect on comparisons between two life-improving interventions, but it drives comparisons between life-improving and life-saving ones. See our linear WELLBY analysis. ↩
- On whether the 0–10 scale can be treated as linear, and on transformations that preserve or reverse conclusions, see Kaiser & Vendrik (above) and Kaiser & Lepinteur, Measuring the unmeasurable. Kimball's argument that scale curvature is partly an ethical choice about inequality aversion is in the workshop transcript. ↩
- Benjamin et al., Guide for calculating benchmarks using calibration questions. Related work on what people actually value, beyond life satisfaction alone: Benjamin et al., What do people want?, and the companion SWB experiments. Full workshop reading list: readings. ↩