Annotate: Use the Hypothes.is browser extension to leave comments on this page. Public annotations are picked up by the workshop Slack bot. For private notes, create a private group.

Wellbeing Workshop · Beliefs Analysis

WELLBY measurement value, conversion, and forecasting · March 16 2026 workshop · The Unjournal · Internal draft, last corrected August 3 2026 — estimates visible, for internal review only

PQ1A: Linear WELLBY Value
Sub-questions
Individual Responses
Methods & Notes

PQ1A: What share of the decision value of the best available wellbeing measure is captured by using a simple linear WELLBY for all interventions? Respondents gave an expected share (0–100%) and an 80% credible interval. This is not a probability that the measure is “reliable enough.”

Dataset filter
Researcher / Academic
Practitioner
Anonymous / unknown
PQ1A expected-value share

Individual expected-value shares with 80% credible intervals · linear scale 0–100%

Each shaded bar spans the respondent's 80% CI. Dot = central estimate. Sorted by central estimate (high to low).

June 18 update: Added Miles Kimball's June 17 response. The most visible effects are a higher forecast for GiveWell-linked charity uptake (new max 50%), a higher forecast for agreement among surveyed economists and practitioners (60%), and a much wider DALY/WELLBY conversion range because Kimball gives 1 WELLBY/DALY with a 0.1–10 interval while stressing that his team has not studied that conversion directly.
Data-quality warning: The form opened with values already set on some sliders — PQ1A at 70 with its 80% interval pre-set to 40 and 90, PQ3A at 25%, PQ3C at 35% — and it recorded only the final value, not whether a respondent moved anything. These were not neutral midpoints: on a 0–100 scale the form presented what looked like a complete, plausible answer, which is a stronger anchor than a midpoint would have been. Exact-default values may therefore be deliberate answers, anchoring effects, or untouched sliders, and the data cannot distinguish them. Among the 7 responses, one matches all three PQ1A defaults, two match the PQ3A default, and two match the PQ3C default.

Resolution (August 10, 2026) — this caveat is now permanent. Each affected respondent was asked directly, twice, which of their values were intended. None answered that question. We are therefore closing it with a stated rule rather than leaving it open: an exact-default value is treated as a non-response. The preferred aggregates below exclude them — PQ1A mean 74.2%, median 78.5% (n=6); PQ3A mean 17.4%, median 10% (n=5); PQ3C mean 36.6%, median 34% (n=5). Including every submitted value instead gives PQ1A mean 73.6%, median 75% (n=7); PQ3A mean 19.6%, median 15% (n=7); PQ3C mean 36.1%, median 35% (n=7). Both sets are shown throughout; the charts plot every submitted value, and nothing has been deleted.

One case cuts against the rule. Kimball's PQ1A sits on all three defaults, but he wrote substantive reasoning on that same question and submitted in June, three months after the workshop, having deliberately returned to the form. That is weak evidence of engagement rather than an untouched slider, and the rule discards it anyway. It is applied uniformly because a rule applied case-by-case on our own reading of who "seems engaged" is worse than a blunt one applied consistently. Readers who disagree can use the all-values figures above.
Interpretation notes:
  • Caspar Kaiser's 100% share is conditional on the comparator being other available measures. He notes that the share would be about 50% (with an interval of roughly 1–90%) relative to the best possible measure. His wide interval [20–100%] reflects this framing uncertainty. Corrected August 3, 2026: this previously read “about 0.5%”, misreading the fraction 0.5 in his form response as a percentage. Correction confirmed by Kaiser.
  • One response shown as "Anon. participant 2" is anonymized at the respondent's request.
  • Dan Benjamin (UCLA), one anonymous respondent, and Miles Kimball (CU Boulder) submitted post-workshop (Mar 17–Jun 17).

PQ1B — Recommended measure for funders

What measure should funders focused on quality-of-life improvements use? All 7 respondents answered.

There is broad agreement that calibration is worth pursuing. The dividing line is whether calibrated WELLBY alone suffices, or whether a broader composite (multiple wellbeing dimensions + revealed preference anchors) is preferable. Plant is the sole advocate for linear WELLBY as the best currently available, without endorsing calibration as a first step.

PQ2 — DALY / WELLBY conversion factor

How many WELLBYs equal 1 DALY? 5 of 7 respondents gave a numeric estimate with an 80% CI. Plant declined on principle; Benjamin left blank.

Plant's objection: "The evidence suggests that for comparing disability weights and wellbeing weights, you get different answers based on the intervention and problem under consideration. So I'd resist doing a simple exchange rate." McGuire similarly advocates separate conversion rates for mental vs. physical health contexts.

PQ3 — Forecast and consensus questions

Point estimates only (no CIs collected for PQ3).

PQ3A — P(>50% of GiveWell top charities include WELLBY-based CEA by 2030)

Preferred (exact-default values treated as non-responses): 5 responses · range 5–50% · mean 17.4% · median 10%. All submitted values: 7 responses · mean 19.6% · median 15%.

PQ3B — Predicted share of surveyed economists/practitioners agreeing linear WELLBY is reasonably useful

6 responses (Anon. participant 2: no answer) · range 15–67% · mean 37% · median 30%

PQ3C — P(calibration changes Founders Pledge's top-five intervention ranking)

Preferred (exact-default values treated as non-responses): 5 responses · range 14–70% · mean 36.6% · median 34%. All submitted values: 7 responses · mean 36.1% · median 35%. “Meaningful change” meant that at least one intervention enters or leaves the top five, or the top-ranked intervention changes.

Click a card to expand. One response anonymized at respondent's request.

Data collection

Responses collected via Netlify form at uj-wellbeing-workshop.netlify.app/beliefs during and after the March 16, 2026 workshop. Raw data stored in wellbeing-beliefs-elicitation-submissions.json (private repo). Two submissions excluded, leaving 7 substantive responses: one test submission, and one (received Mar 17) whose entire content was the form's five slider default values with no name, email, or free text of any kind — someone opened the form and submitted without answering. It was previously shown here as “Anon. participant 1”. Presenting it as a participant's beliefs was a mistake on our part, and including it pulled every statistic toward the defaults. Removing it moves the PQ1A median from 72.5% to 75% and the PQ3A median from 20% to 15%; means barely change.

Response timing

Credible intervals

The form asked for a central estimate and an 80% credible interval (10th to 90th percentile) for PQ1A and PQ2. No CIs were collected for PQ3A, PQ3B, or PQ3C (point estimates only). Not all respondents provided CIs for all questions.

Visible slider defaults

PQ1A and two PQ3 questions used visible, submittable defaults, and the form did not record whether a respondent touched a slider. Exact-default entries therefore cannot be distinguished from deliberate responses. The affected defaults were PQ1A = 70/40/90, PQ3A = 25%, and PQ3C = 35%. Future workshop forms now use blank, click-to-place sliders and omit untouched values.

Each affected respondent was asked directly which of their values were intended — on July 30 and again on August 3. None answered. Kimball replied but returned the question rather than answering it (“If it shows as the default, does that suggest I didn't answer it at all?”), was given a specific answer and a restatement of which two of his values were affected, and did not come back on it. Kaiser answered on other points but not this one. The anonymised respondent did not reply at all.

Closed August 10, 2026 with a stated rule: an exact-default value is treated as a non-response. The preferred aggregates exclude them; figures on all submitted values are reported alongside throughout, and no value has been deleted from the data or the charts. This is not repairable retrospectively, so it is recorded as a permanent limitation of this elicitation rather than an open question. The form itself was fixed after the workshop: sliders now start blank and untouched values are not submitted, so later workshops do not carry this problem.

Correction round and publication consent

Each named respondent was emailed individually on July 30–31, 2026, asked to check their own card and to confirm whether it may be published under their name, with a reply deadline of August 6. Status as of August 10:

RespondentCard confirmedName may be publishedNotes
Dan BenjaminYesYesAlso confirmed the standardised-improvement attribution used in the write-up
Caspar KaiserYesYesOne correction, applied: the “possible measures” figure is 50%, not 0.5%
Miles KimballYesYesAsked again August 3 which of his values were intended; no answer yet — see above
Julian JamisonYesYes, fully publicConfirmed August 3, after checking that others were going public under their own names
Michael PlantYesYesConfirmed August 10; declined post-hoc changes on principle, having seen others' results
Joel McGuireYesYes (implicit)Replied substantively August 10, confirming the conversion argument quoted in the write-up; did not object to naming after two explicit opt-out offers
Anon. participant 2Remains anonymisedContacted July 30 and again August 3 about anonymity and the slider defaults; no reply, so the anonymised treatment stands unchanged

All six named respondents have now confirmed their cards and agreed to publication under their own names (McGuire's consent is implicit, as noted above). The anonymised respondent was told that silence keeps the anonymised treatment unchanged, and has not replied, so it stands. The correction round is closed as of August 10.

The consent conditions for wider sharing are met. The page remains unlisted until a deliberate publication decision is made; the remaining editorial caveat is the unresolved slider-default question above (Kimball and Kaiser), which affects how the aggregates should be presented, not whether cards may be shown.

Anonymization

One respondent (shown as "Anon. participant 2," affiliation: funder / evaluator) requested that their individual response not be shared publicly. Their quantitative estimates are included in aggregates and individual CI charts. Their qualitative reasoning is shown in summary form in their response card. This page uses an unlinked internal URL and is not indexed from the main workshop site.

HLI concentration note

3 of the 5 workshop-day named respondents have HLI affiliations (Plant: founder/director; Kaiser: board chair; McGuire: researcher). The Benjamin, Jamison, and Kimball responses provide non-HLI academic perspectives; Benjamin and Kimball are co-authors of related scale-use and multi-dimensional wellbeing work. Interpret aggregate figures accordingly — they do not represent a balanced cross-section of the field.

Completeness

Question N responses Notes
PQ1A central estimate8All 8 provided
PQ1A 80% CI8All 8 have submitted values; two exactly match the visible default triple
PQ1B measure recommendation7All respondents answered
PQ2 DALY/WELLBY estimate + CI5Plant declined (principled); Benjamin left blank
PQ3A GiveWell-linked charity uptake by 20308Three values exactly match the visible 25% default
PQ3B expected expert/practitioner agreement6Anon. participant 2 (ran out of time) did not answer
PQ3C calibration changes FP top five8Three values exactly match the visible 35% default