We measure AIs to see whether they can pass a bar exam, write working code, and use your computer interface. We test to see how good they are at completing complex tasks, or just impressing humans.
Thank you for pointing this out. Western, liberal, English-speaking, global north frameworks are embedded in so much, shaping the metrics, data collection, and standarization that undergirds so much of our technology. I’ll defintely be heading over to those posts (The English Machine, Indigenous Machines).
The social media parallel proves the very point you're raising — by the time measurement was strong enough to act on, the harm had already proliferated. Measurement is a lagging indicator.
The harder question is what a leading indicator would even look like. I'd argue it's not observable from outside the person at all — it's whether the capacity for genuine inner inquiry is developing or eroding. That question goes deeper than sycophancy scores or dependency metrics.
I agree. Measurement is a lagging indicator. We already know that the dependency is there. The recent studies on AI companion use amongst teenagers indicate harm is well and truly here.
I also agree that the more fundamental question and that is much harder to measure is whether we as humans are genuinely able to critically engage with AI in our use (not be led). Perhaps one way of doing this would be simulating our ability to distinguish fact from fiction (although this problem exists also with or without AI).
Critical engagement with AI isn't a skill you can train directly. Research from Harvard and Max Planck found that thinking more carefully about AI output actually reduces evaluation accuracy, not improves it. The cognitive layer is the wrong instrument for this problem. What's needed is the capacity that operates before analysis begins, the ability to sense when something is off before you can explain why. When that instrument is intact, AI's confident tone doesn't override what was already registered as wrong. When it's degraded, the analysis arrives too late or not at all. That's what's worth developing.
"Measurement is a lagging indicator.” Yup. We need to be preemptive in safety and regulation, but that isn’t a tune that Silicon Valley and Big Tech (or America for that matter) likes to listen to.
Regulation will always arrive after the harm. The more reliable protection starts at home. Parents who develop their own capacity to notice what's working and what isn't pass that to their children. Not as a rule. As a way of being. That transmission happens long before any regulation does.
Loved this piece. Loved this framing. Sycophancy, dependency, emotional reliance, critical thinking, trust - these things don’t always show up in a single prompt and response. They unfold over time, in relationship, in repetition, in the small ways people start adapting themselves around a system. This is a really important topic, and one I keep circling in my own writing too - not just whether a tool works, but what kinds of habits, dependencies, shortcuts, and risks start forming around it once people bring it into ordinary life. Thank you for sharing :)
Integrate a BioPsychoSocial evaluation linked to a functional brain health approach, and ai impacts can be subject to metrics which track impacts where humans live.
We have detailed frameworks for what AI systems should do — accuracy benchmarks, bias audits, performance metrics. We have almost nothing for what they do to us: attention fragmentation, reduced tolerance for ambiguity, atrophying of skills we've offloaded, changes in how we reason when we know a model will finish our sentences.
EU AI Act requires impact assessments for high-risk systems. But "high-risk" is defined by sector and application type, not by cognitive or social consequence. A recommendation system that gradually reshapes how millions of people form opinions isn't classified as high-risk. A healthcare diagnostic tool is.
The question "why aren't we measuring it" has a structural answer: there's no regulatory incentive to, and the harms are diffuse, delayed, and hard to attribute to any single system. That's exactly when governance should step in — and currently doesn't.
Because measuring it is not profitable, at least not in the short term.
We saw the same pattern with social media. The harm to attention, mental health, social trust, and the human nervous system became increasingly visible, but the business model rewarded engagement, not well-being.
So the damage was treated as externalities.
AI may follow the same path, but with even deeper consequences. It is not only shaping what people see. It is shaping how they think, write, learn, decide, and relate to their own judgment.
That matters especially for young people and students, who are still building the cognitive muscles AI can so easily bypass.
If we only measure adoption, usage, and productivity, we may miss the most important question:
Yes. And the gap isn't only psychological. One thing AI does to us is decide who it serves and who it prices out: the same sentence can cost up to 10x more to run in some languages than in English, and those speakers are often understood worst. What's striking is that this harm is already cleanly measurable. So part of the gap isn't that the tools are immature.
It's that we're not looking even where we easily could.
This question hits hard. We have years of research on what social media does to attention spans and mental health, but almost nothing systematic on what AI interactions are doing to us at scale. The lack of measurement isn't accidental either. Worth pushing hard on this.
One area I'm especially drawn to is what happens to understanding itself. We have many ways of measuring what AI can do, but it's much harder to measure what happens when people begin outsourcing parts of the process through which understanding forms. I recently explored that question in an essay called The Messy Middle, if you're interested. https://www.heavenandearth.com/p/the-messy-middle-thinking-with-ai?r=1lcuyw
I am not a researcher. I am a computer science person turned human systems coach. And I wonder if things like how satisfaction, engagement, and sense of ownership change with with how and how much people are cognitively offloading to AI would be a helpful to measure. It is not just AI that is the problem, but how people are using it. And how work environments and cultures are adapting (or not) to help employees have the breathing space/thinking space/downtime during the day to reset and recharge to avoid "brain-fry" and decision fatigue. Might measuring the human effects of using AI be helpful in addition to measuring the AI itself?
Guys, what do you mean "we don't measure it?" What are you even talking about?? This is one of THE hottest topics there is right now in cog/neuro-science!
Underneath the three problems is a fourth: the constructs themselves are
English-shaped.
"Dependency", "critical thinking", "validated ideation" are
anglophone-psychology categories. Benchmarks are written in English,
evaluated by English-speaking judges, scored against rubrics built in
English. A harm that doesn't surface as one of those constructs doesn't
get measured.
Languages with relational grammar — Basque allocutive, Quechua
evidentials, Indigenous Australian kinship — embed psychosocial
relations English speakers either spell out longhand or do without.
When speakers of those languages route their communication through a
model trained ~90% on English, what they lose compounds across generations
rather than within sessions — invisible to single-turn and multi-turn
evaluation alike.
Independence, alignment, shared methodology — necessary, but not
sufficient if the field defines its constructs in the same language
whose limits it's trying to evaluate. I've been working this through
on Substack (The English Machine, Indigenous Machines), with a Basque
piece next that walks through where even four billion tokens of
fine-tuning hits a ceiling.
Thank you for pointing this out. Western, liberal, English-speaking, global north frameworks are embedded in so much, shaping the metrics, data collection, and standarization that undergirds so much of our technology. I’ll defintely be heading over to those posts (The English Machine, Indigenous Machines).
The social media parallel proves the very point you're raising — by the time measurement was strong enough to act on, the harm had already proliferated. Measurement is a lagging indicator.
The harder question is what a leading indicator would even look like. I'd argue it's not observable from outside the person at all — it's whether the capacity for genuine inner inquiry is developing or eroding. That question goes deeper than sycophancy scores or dependency metrics.
Explored this from a different angle here: https://newsletter.awarelife.co.il/p/the-ai-revolution-is-not-about-technology
I agree. Measurement is a lagging indicator. We already know that the dependency is there. The recent studies on AI companion use amongst teenagers indicate harm is well and truly here.
I also agree that the more fundamental question and that is much harder to measure is whether we as humans are genuinely able to critically engage with AI in our use (not be led). Perhaps one way of doing this would be simulating our ability to distinguish fact from fiction (although this problem exists also with or without AI).
Critical engagement with AI isn't a skill you can train directly. Research from Harvard and Max Planck found that thinking more carefully about AI output actually reduces evaluation accuracy, not improves it. The cognitive layer is the wrong instrument for this problem. What's needed is the capacity that operates before analysis begins, the ability to sense when something is off before you can explain why. When that instrument is intact, AI's confident tone doesn't override what was already registered as wrong. When it's degraded, the analysis arrives too late or not at all. That's what's worth developing.
"Measurement is a lagging indicator.” Yup. We need to be preemptive in safety and regulation, but that isn’t a tune that Silicon Valley and Big Tech (or America for that matter) likes to listen to.
Regulation will always arrive after the harm. The more reliable protection starts at home. Parents who develop their own capacity to notice what's working and what isn't pass that to their children. Not as a rule. As a way of being. That transmission happens long before any regulation does.
Important diagnosis.
The harder question is sequence. Measurement comes after formation.
By the time psychosocial evaluations can detect impact, the conditions that shape the user's thinking have already shifted.
You can build instruments for what's downstream of formation. The formation itself sits at a layer that doesn't show up under any instrument.
The frame may need to be wider than evaluation.
Loved this piece. Loved this framing. Sycophancy, dependency, emotional reliance, critical thinking, trust - these things don’t always show up in a single prompt and response. They unfold over time, in relationship, in repetition, in the small ways people start adapting themselves around a system. This is a really important topic, and one I keep circling in my own writing too - not just whether a tool works, but what kinds of habits, dependencies, shortcuts, and risks start forming around it once people bring it into ordinary life. Thank you for sharing :)
I worry about a time when we'll start asking AI "What do I think of this?" relying on it for our beliefs and values.
Integrate a BioPsychoSocial evaluation linked to a functional brain health approach, and ai impacts can be subject to metrics which track impacts where humans live.
The measurement gap is also a governance gap.
We have detailed frameworks for what AI systems should do — accuracy benchmarks, bias audits, performance metrics. We have almost nothing for what they do to us: attention fragmentation, reduced tolerance for ambiguity, atrophying of skills we've offloaded, changes in how we reason when we know a model will finish our sentences.
EU AI Act requires impact assessments for high-risk systems. But "high-risk" is defined by sector and application type, not by cognitive or social consequence. A recommendation system that gradually reshapes how millions of people form opinions isn't classified as high-risk. A healthcare diagnostic tool is.
The question "why aren't we measuring it" has a structural answer: there's no regulatory incentive to, and the harms are diffuse, delayed, and hard to attribute to any single system. That's exactly when governance should step in — and currently doesn't.
Because measuring it is not profitable, at least not in the short term.
We saw the same pattern with social media. The harm to attention, mental health, social trust, and the human nervous system became increasingly visible, but the business model rewarded engagement, not well-being.
So the damage was treated as externalities.
AI may follow the same path, but with even deeper consequences. It is not only shaping what people see. It is shaping how they think, write, learn, decide, and relate to their own judgment.
That matters especially for young people and students, who are still building the cognitive muscles AI can so easily bypass.
If we only measure adoption, usage, and productivity, we may miss the most important question:
What is this doing to the human mind?
I feel seen here. I just wrote a post (less academic) about this erosion of our ability to handle friction https://substack.com/home/post/p-190418093
this feels like the social media mistake all over again
Yes. And the gap isn't only psychological. One thing AI does to us is decide who it serves and who it prices out: the same sentence can cost up to 10x more to run in some languages than in English, and those speakers are often understood worst. What's striking is that this harm is already cleanly measurable. So part of the gap isn't that the tools are immature.
It's that we're not looking even where we easily could.
This question hits hard. We have years of research on what social media does to attention spans and mental health, but almost nothing systematic on what AI interactions are doing to us at scale. The lack of measurement isn't accidental either. Worth pushing hard on this.
One area I'm especially drawn to is what happens to understanding itself. We have many ways of measuring what AI can do, but it's much harder to measure what happens when people begin outsourcing parts of the process through which understanding forms. I recently explored that question in an essay called The Messy Middle, if you're interested. https://www.heavenandearth.com/p/the-messy-middle-thinking-with-ai?r=1lcuyw
I am not a researcher. I am a computer science person turned human systems coach. And I wonder if things like how satisfaction, engagement, and sense of ownership change with with how and how much people are cognitively offloading to AI would be a helpful to measure. It is not just AI that is the problem, but how people are using it. And how work environments and cultures are adapting (or not) to help employees have the breathing space/thinking space/downtime during the day to reset and recharge to avoid "brain-fry" and decision fatigue. Might measuring the human effects of using AI be helpful in addition to measuring the AI itself?
The problem is that we still measure AI like a tool.
Accuracy. Speed. Benchmark scores. Productivity gains.
But the real shift is architectural.
AI is starting to rewire how humans externalize memory, judgment, recursion, attention, and even identity itself.
We are not just outsourcing labor anymore.
We are outsourcing parts of cognition.
The dangerous part is not “AI replaces humans.”
It’s that humans slowly stop building internal structure because external cognition becomes cheaper than internal cognition.
Civilization may become more intelligent while individual humans become cognitively thinner.
That paradox is still massively under-measured.
Guys, what do you mean "we don't measure it?" What are you even talking about?? This is one of THE hottest topics there is right now in cog/neuro-science!
https://thealgorithmicbridge.substack.com/p/what-the-studies-say-about-how-ai