What Gets Measured, Gets Built
AI is optimized for what it can do. A new program asks what it does to us. The case for humane evals.
Every AI model now ships with a scorecard: how well it codes, reasons, passes the bar exam. We have an entire industry devoted to measuring what these systems can do. We have almost nothing devoted to measuring what they do to us. When the only numbers that exist are capability numbers, those are the numbers everyone competes to win.
There is a quiet law inside every technology company: what gets measured gets optimized [1]. Right now the entire apparatus of AI development is pointed at measuring how capable and powerful models are, and then narrowly optimizing for exactly those metrics while the downstream consequences go uncounted [1]. This is not a moral failing so much as a structural one. You cannot manage, reward, or race toward a thing you have no instrument for. And the instruments we have all point in one direction.
The Center for Humane Technology's answer is a program called Humane Evals, and its premise is a simple inversion: instead of measuring what AI can do, start measuring what AI does to us [1]. The ambition goes further than harm-avoidance. The stated aim is to flip the competitive dynamic itself, so that instead of races to the bottom on capability and engagement, you could incentivize races to the top on safety, or better still on making people more resilient and more developed human beings [1]. Put plainly: build a scoreboard for human flourishing, and companies will start playing to it, because companies play to whatever scoreboard exists.
Why does this matter for attention specifically? Because the effects that engagement optimization produces are exactly the kind that current metrics are blind to. They are slow, they are distributed across a population, and they show up not as a product failure but as a change in the person using the product. Consider what the clinical literature has been quietly documenting. In a study of 491 South African university students, heavier social media use was linked to poorer mental health, with smartphone addiction acting as the pathway connecting the two rather than a mere correlate [3]. The distress does not announce itself the moment you open an app. It accumulates through a pattern of use that a capability benchmark would never register.
The mechanism gets more specific in a structural model of Jordanian students, where fear of missing out and social phobia predicted smartphone addiction, but academic procrastination sat in the middle as a mediating step [4]. In other words, the anxiety does not go straight to the phone; it routes through avoidance, and the device becomes the instrument of that avoidance. This is the sort of causal chain that a company optimizing for daily active minutes has no reason to see and every reason not to look for. The engagement metric is satisfied either way.
Here is where the two threads meet. The reason measuring what AI does to us is hard is the same reason it is necessary. The damage is downstream, mediated, and easy to attribute to the user rather than the system. A person who cannot sleep, cannot focus, or feels worse after an hour online reads it as a personal weakness, and the product's scorecard agrees, because the product's scorecard only knows whether they came back. Humane Evals is an attempt to build the missing measurement so the harm stops being invisible [1].
The honest caveat: none of this is settled, and a scoreboard is only as good as what it counts. The clinical evidence on whether we can even treat problematic digital use remains fragmented across very different behaviors, which is precisely why researchers keep calling for systematic comparison [2]. Measuring flourishing is harder than measuring throughput, and a poorly designed humane metric could be gamed as easily as an engagement metric. But the alternative is the status quo, where the only questions we can answer with numbers are the ones that happen to serve the people selling the product. The first act of reclaiming attention may be insisting that the effect on attention gets counted at all.
RESEARCH RADAR
- A humane scoreboard. CHT's Humane Evals program aims to measure what AI does to people rather than only what it can do, on the logic that whatever gets measured is what gets optimized [1]. The goal is to redirect competition from engagement toward resilience and human development [1].
- Addiction as the pathway. Among 491 South African university students, smartphone addiction functioned as the mechanism linking heavier social media use to psychological distress, not just a companion symptom [3]. The pattern of use, not the mere fact of use, carried the harm.
- Anxiety routes through avoidance. In a structural model of Jordanian students, procrastination mediated the link between fear of missing out, social phobia, and smartphone addiction [4]. The phone is often the endpoint of an avoidance loop that starts with anxiety.
ONE THING TO TRY
Pick one app you open reflexively and ask a different question than "how long was I on it." Ask: did I feel better or worse after, and what was I avoiding when I reached for it. Write the one-word answer down. That is the metric no dashboard gives you, and it is the one that matters.
WORTH YOUR ATTENTION
- We Measure What AI Can Do. We Should Measure What It Does to Us. (Your Undivided Attention) — The clearest articulation of why the missing scoreboard is the whole problem [1].
- The Negative Mental Health Consequences of Social Media Use in South Africa (Behavioral Sciences) — Careful evidence that addiction, not screen time, is the pathway to distress [3].
- Academic Procrastination as a Mediator (BMC Psychology) — Maps the anxiety-to-avoidance-to-phone chain that engagement metrics never see [4].
- Therapeutic Interventions Targeted at Problematic Use of Digital Technology (JMIR Mental Health) — A systematic review honest about how fragmented the treatment evidence still is [2].
We opened with a scorecard that only counts capability. The deeper point is that a scoreboard is never neutral: it decides what gets built. If what gets measured gets optimized [1], then the most consequential thing we can do for our attention is not to use less, but to insist the world start counting what these tools cost us — and start rewarding what they give back.
Sources
- [1] We Measure What AI Can Do. We Should Measure What It Does to Us.: Your Undivided Attention (Center for Humane Technology)
- [2] Therapeutic Interventions Targeted at Problematic Use of Digital Technology: Systematic Review and Meta-Analysis of Evidence.: JMIR mental health
- [3] The Negative Mental Health Consequences of Social Media Use in South Africa: The Role of Smartphone Addiction.: Behavioral sciences (Basel, Switzerland)
- [4] Academic procrastination as a mediator linking fear of missing out and social phobia to smartphone addiction among university students: a structural model.: BMC psychology