Thinking in AI
“This invention will produce forgetfulness in the minds of those who learn to use it. They will appear wise, but not be so.” — Thamus to the god Theuth, on the invention of writing, in Plato’s Phaedrus
Sometime in the last year I caught myself doing something I did not like. I was reading a paper, a genuinely hard one, and about four paragraphs in I felt a familiar pull toward the tab where a model sat waiting. Not to check a citation. To be told what the paper said. I did not act on it. But I noticed the pull, I noticed that it arrived earlier than it used to, and I noticed the exact moment it arrived: the moment the reading got difficult.
That is one small observation from one aging data wrangler, and I want to be careful about the weight it can carry. So I went looking for what the research says, expecting the usual moral panic dressed in citations. I found something stranger.
The story we have been telling is that AI is making us stupid. The research does not say that. What it says is stranger.
What the Electrodes Saw
Start with the study everyone quotes and almost nobody reads. Researchers at the MIT Media Lab put fifty-four people in EEG caps and had them write essays across several sessions. One group used ChatGPT. One used a search engine. One used nothing but the contents of their own skull.
The brain-only group showed the strongest and most widely distributed neural connectivity. The search group landed in the middle. The ChatGPT group showed the weakest connectivity, concentrated in the networks tied to memory, attention, and analysis. They also struggled to recall what they had written minutes earlier, and their prose came out generic. The authors named the effect cognitive debt, which is a good name: the tool does the synthesis, the brain stays quiet, and the balance comes due later.
Here is the finding that got buried under the headlines. Order of use mattered enormously. Participants who wrote unaided first and got the model later showed stronger recall and broader brain activity than participants who started with the model and were then made to work alone. The same tool, the same task, the same people. Reverse the sequence and you reverse the sign.
The tool helped when it scaffolded thinking already underway. It hurt when it substituted for thinking that had not started yet.
One more thing, because my regular readers know I try not to oversell. That study is a preprint, its sample is small, and one of the MIT researchers said plainly, “We didn’t find any brain rot.” Half the internet reported that they had. Hold the finding loosely and the sequencing insight tightly, because everything else here turns on it.
Twenty Percent Faster
Now the result that changed how I think about all of this.
METR ran a randomized controlled trial with sixteen experienced open-source developers working two hundred forty-six real tasks on repositories they personally maintain. Real code, real stakes, their own projects. Each task was randomly assigned to permit AI tools or forbid them.
The developers were nineteen percent slower with AI.
Before the trial, they predicted AI would make them twenty-four percent faster. Afterward, having personally lived through every one of those slower tasks, they estimated they had been twenty percent faster. The gap between what they felt and what the clock recorded survived direct contact with the clock.
Read that again, because it is the whole argument. These were experts working in codebases they knew intimately, and they could not perceive a nineteen-point productivity loss while it was happening to them.
Set that beside two other findings and a shape emerges. Microsoft Research and Carnegie Mellon surveyed three hundred nineteen knowledge workers across nine hundred thirty-six real AI-assisted tasks and found a seesaw: the more confidence a worker placed in the AI, the less critical thinking they reported doing, while the more confidence they placed in their own skill, the more they did. And a 2026 study with the deadpan title “Confidence Without Competence in AI-Assisted Knowledge Work” found that the plain chatbot interface produced the highest perceived understanding alongside the lowest measured learning of any condition tested.
So the damage is not to our thinking. It is to the instrument that tells us whether we are thinking at all.
Effort Blindness
I have come to call this Effort Blindness, and I propose it as a companion to the Blind Spot Multiplication Problem I wrote about in Situational Blind-Spots. A blind spot hides a thing in the world. Effort Blindness hides your own cognitive state from you. It is the more dangerous of the two, because every correction mechanism we have depends on a working sense of how hard we are working.
Consider the shape of the failure. When a task feels effortful we slow down, check ourselves, seek a second opinion. That feeling triggers every metacognitive habit we own. Fluent output removes the feeling without removing the difficulty, which sits in the work unexamined because the alarm that should have fired never did.
Harvard and BCG researchers found something adjacent when they gave GPT-4 to seven hundred fifty-eight consultants. Inside the frontier of what the model does well, large gains. Outside it, AI users performed nineteen percentage points worse than the control group. They called that boundary the jagged frontier, and jagged is the word that matters, because a jagged edge cannot be seen from inside the task. You cannot feel which side of it you are standing on.
What the Endoscopists Forgot
Abstractions are cheap. Here is what this costs in a room where it counts.
Four endoscopy centres in Poland introduced an AI assistant for polyp detection during colonoscopy. Researchers compared three months before to three months after, and published the result in The Lancet Gastroenterology & Hepatology. The metric is the adenoma detection rate, which is how often a doctor finds the precancerous growths that are the entire point of the procedure.
Unassisted detection fell from 28.4 percent to 22.4 percent.
Six percentage points of diagnostic skill, gone, in twelve weeks. These were practicing physicians in a high-stakes specialty, and the researchers believe it is the first documented deskilling effect from clinical AI. Nobody decided to get worse at looking. They simply stopped doing the looking themselves, and the capacity went where unused capacities go.
The same pattern shows up in a cleaner experiment with a thousand Turkish high school students learning mathematics. One group got plain GPT-4. One got a version wrapped in teacher-designed hints and safeguards. One got nothing. With the tools in hand, practice performance jumped forty-eight percent. Then the researchers took the tools away and gave everyone an exam.
The plain-GPT group scored seventeen percent worse than students who never had access at all. Worse than nothing. They had used the model as a crutch, and crutches do to legs what crutches do.
The safeguarded group came through nearly unharmed. Same model, same students, same subject. The difference was entirely in how the interaction was designed, which tells you the tool was never the variable we should have been arguing about.
The Center of the Distribution
Individual atrophy is the part of this we can measure. There is a collective effect that no individual can feel at all, and it worries me more.
When researchers gave writers AI-generated plot ideas, individual stories got more creative. Measured across all the writers, the stories converged. Each author gained; the group’s variety shrank. A 2026 analysis of nearly seven thousand essays found the same thing and gave it a name: a quality-homogenization tradeoff, real gains in each piece bought with real narrowing across the corpus.
The mechanism is a loop, and it is the kind of loop I have spent my career watching systems fall into. A language model generates toward the center of its training distribution. Widespread use means most humans now encounter that center most of the time. Human writing shaped by that center gets scraped and becomes training data. The distribution narrows. Repeat.
Nobody chose this. No committee approved it. It is an emergent property of a few hundred million people independently reaching for the most fluent available sentence, and it runs whether or not anyone intends it. Buckminster Fuller taught us to do more with less. He did not warn us that doing more with less might mean all of us doing the same more, with the same less.
Then there is the flattery problem. In the spring of 2026 teams at MIT and Stanford published on chatbot sycophancy, including a formal proof and a preregistered study in Science. Brief conversations with agreeable models left people holding more extreme and more certain beliefs, and enjoying the conversation more. The models mirror a user’s politics more precisely the better they can infer them.
Someone at Futurism called them Dunning-Kruger machines, and I have not been able to improve on the phrase. A system that raises your certainty while lowering your accuracy has stopped being a research assistant and become a mirror with opinions. Me, I am often wrong but never in doubt. That joke seems to becoming pervasive.
Boxtown Is Doing the Reading
Now walk outside the essay and look at what people are actually doing.
In 2024 xAI moved into a shuttered Electrolux plant in the Boxtown neighborhood of South Memphis and stood up a supercomputer called Colossus. Members of the public, not the regulator, discovered it was being powered by thirty-five gas turbines with no permits. For Colossus 2 across the state line in Southaven, Mississippi, the NAACP sued in April of 2026 over twenty-seven more unpermitted turbines, represented by the Southern Environmental Law Center and Earthjustice. Whitehaven, Boxtown, and Riverside Park sit in census tracts scoring between the ninety-fifth and ninety-ninth percentile nationally on the EPA’s own cumulative pollution burden screen. The people breathing that air already carried the asthma and cancer rates to prove it.
Memphis is the sharp end. The broad end is arriving on everyone’s utility bill. PJM, the largest grid operator in the country, attributes a 6.3 billion dollar increase in consumer electricity costs over three years mostly to data center demand. Electricity prices rose 6.9 percent in 2025 against headline inflation of 2.9 percent. Goldman projects data centers going from 4.1 percent of peak summer demand to 8.5 percent within two years.
People noticed. In the first quarter of 2026 alone, roughly seventy-five projects worth about a hundred thirty billion dollars were blocked or delayed, matching the entire preceding year in three months. Organized opposition groups went from three hundred ninety-six at the end of 2025 to eight hundred thirty-three by March, spread across forty-nine states.
Here is what I want you to notice, and it is the reason this section sits in an essay about cognition.
Look at what those eight hundred thirty-three groups actually do. They read county comprehensive plans. They file Right-to-Know requests and wait for the records. They learn the difference between a special-use permit and a rezoning, track water and utility reviews, master the variance criteria well enough to prove a project fails them, and turn out enough neighbors that a commission has to count the room at seven in the evening on a Tuesday. Data Center Watch, which tracks this for the industry, wrote that communities have “internalized an opposition playbook.” That is hard practical knowledge, developed under adversarial conditions by people who mostly are not lawyers, and transmitted across forty-nine states in months.
None of it can be offloaded. There is no prompt that reads the comprehensive plan for you, and there is certainly no prompt that shows up at the hearing.
So the most sustained unassisted collective reasoning in America right now is being done by ordinary people organizing against the physical infrastructure of artificial intelligence. The civic muscle we keep eulogizing is being worked hardest by the people with the most immediate reason to distrust the machine, and they are winning often enough to have bent the national buildout.
I find that clarifying. Cognitive capacity did not disappear. It relocated to wherever the stakes were concrete enough to make thinking feel worth the effort. Which returns us, uncomfortably, to effort.
Watch what happened next, because it is the best illustration of Effort Blindness I can offer, and it happened to the people least likely to admit susceptibility.
As the blockade mounted, a theory took hold among wealthy technologists: the opposition is a Chinese influence operation. There is a real fact underneath it. OpenAI reported in June of 2026 that it had banned a cluster of accounts, which it named Data Center Bandwagon, run from China through VPNs, posing as Americans, generating comic strips blaming AI data centers for household electricity bills. Chinese state outlets have pushed the same line. All true, all documented.
Two things about that finding rarely travel with it. OpenAI also reported that the operation produced virtually no authentic engagement and appeared to have had little or no effect. And the larger claim, that the eight hundred thirty-three domestic groups are foreign-funded, rests on a report from a crypto advocacy nonprofit. NPR went looking. Darren Linvill of Clemson’s Media Forensics Hub, whose profession is finding exactly this kind of campaign, said plainly that they had not found much. The named groups deny it.
So: some of the most sophisticated people alive, handed an explanation that flattered them and cost them nothing, took it at speed rather than engage the objection actually in front of them, which is that electricity prices rose 6.9 percent against 2.9 percent inflation and somebody noticed. The fluent answer arrived. The alarm never fired. That is Effort Blindness viewed from the top of the income distribution.
The Other Half
I have been here before, and I want to tell you exactly when.
In 2001, before blogs were a thing, I published a set of essays called Zen and Java on my own site. The relevant one is reproduced in full at the end of this post, since no copy of it survives anywhere else. One of them proposed something I named Goff’s Law: by any measure of intelligence, half of humanity is and will always be below median IQ. The observation is definitional and I said so at the time. Half of any distribution lies below its middle. That is arithmetic wearing a lab coat.
The germane part came after. If earning capacity tracks measured intelligence, and in a knowledge economy it plausibly does, then what I called the fitness landscape favors the high end, and unless we did something the smart would inherit the earth. Then I aimed the obligation at my own profession. The code we write, I argued, must serve all, not just the wealthy and not just the smart. All.
Twenty-five years later the scorecard is in, and it satisfies nobody.
The obligation is being met, more or less by accident. That study of 5,172 support agents found the gains from AI landing almost entirely on the least experienced and least skilled workers, while the most experienced gained a little speed and gave back a little quality. For once a technology arrived and served the other half first. Nobody designed that outcome as a matter of virtue. It fell out of where the headroom was.
And yet the median is moving. Norwegian military conscript scores peaked with the 1975 birth cohort and have fallen about two tenths of a point a year ever since, and the decline shows up within families, later-born brothers scoring below earlier-born ones, which rules out genetic drift and immigration and leaves the environment holding the bag.
So the floor rises while the ground beneath it settles. The tool helps most precisely those whose measured capability is slipping, and it arrives with Effort Blindness bundled in, which means the half that gains the most also loses the most reliable way to tell whether it gained anything.
The 2001 essay had one more move, and I did not expect it to come back the way it has. Citing the World Watch Institute, I noted that most planetary ecological damage traces to the wealthiest twenty percent and the poorest twenty percent, both ends at once, and that since ecologies bind everyone it does not finally matter how smart or how rich you are if you cannot breathe the air or drink the water.
I meant that as a general principle. In Boxtown it is a description. The wealthiest twenty percent built the compute. The poorest twenty percent are breathing the exhaust from thirty-five turbines nobody permitted. My old argument came back wearing a gas turbine, and I would rather it had stayed a metaphor.
The Shadow Side of the Argument
There is a shadow side here, and my regular readers know I cannot look away from the shadow. This time the shadow falls across my own argument.
Start with the timeline. That Norwegian peak was the 1975 birth cohort. SAT and ACT averages have slid for more than a decade. NAEP showed reading and mathematics declines before the pandemic and steeper ones after. Whatever is happening to human cognition, it began roughly thirty years before anyone typed a prompt, and I was writing about it in 2001 with no chatbot in sight to blame.
Which means the atrophy story flatters us. It hands us a villain with a logo for a curve that started bending when the villain was still a research paper. I distrust any explanation that arrives that conveniently, and you should distrust mine.
The honest ledger has real entries on the other side. Randomized trials of AI tutoring show learning gains that Brookings compares to one or two additional years of conventional schooling, in wealthy and poor countries alike. The safeguarded tutor in the Turkish study did no harm at all. Sequence and design decided every one of these outcomes, which is the least fashionable finding in the whole literature and the most useful.
The protests carry their own shadow too. Some fraction of that opposition is ordinary NIMBYism wearing an environmental costume, and the same procedural machinery that stops a data center has spent fifty years stopping housing, transit, and transmission lines that communities badly needed. Civic capacity is a capability, not a virtue. It can be exercised well or badly, and organizing skill says nothing about whether the organizers are right.
The economics embarrass both camps. Georgia Tech researchers ran ninety-three counties that got their first large data center against roughly three thousand controls and found real gains: private employment up four to five percent, construction up eleven, wages up three to four. Virginia’s legislative auditors credit data centers with 9.1 billion dollars of annual GDP. The industry is telling the truth about that, and quieter about the rest.
Those gains concentrate in metropolitan counties, where employment rose about 4.1 percent and wages about 5.5. In less populous counties the same authors found the spillovers negligible. Virginia’s auditors also project as much as 18 billion dollars in added generation and transmission costs by 2040, against a tax exemption that ate nearly eighty percent of all state economic incentive spending in fiscal 2024. The benefit is real and lands in the metros. The bill is real and lands on everyone. And the fights are concentrated in the rural counties where the promised prosperity never shows up in the data.
And Thamus, for all his rhetorical victory in the Phaedrus, was half wrong. Writing did degrade trained memory, exactly as he predicted. It also built, over centuries, cognitive practices that unaided memory could never have reached, including the dialogue in which he makes the complaint. Socrates’ objection was that a text cannot answer back. Our situation has inverted his: the text answers back beautifully, tirelessly, and far too agreeably.
Let Us Reason Together
I am not going to tell you to stop using these tools. I use them daily, this essay was researched with them, and the fluency they offer is real. What I will insist on is that fluency and understanding have come apart, and that anyone who cannot tell them apart in the moment has already lost the thing worth protecting.
Three practices, and all three follow from the research rather than from my preferences.
-
Sequence deliberately. The MIT result was unambiguous about order: think first, then consult. Write the bad draft, form the wrong hypothesis, attempt the proof, and only then open the model. The machine is excellent at improving a thought that exists and corrosive when it supplies one that does not.
-
Design the interaction rather than accepting the default. Plain chat produced the highest confidence and the lowest learning; the safeguarded tutor eliminated nearly all the harm. Ask the model to argue against you. Ask it what it left out and what would falsify its answer. A system that only agrees is a system giving you nothing you did not bring.
-
Calibrate on purpose, and on a schedule. This one is mine, and it follows from Effort Blindness directly. If the gauge that measures your own effort has been quietly disconnected, no amount of introspection will reconnect it. Take a real task, do it entirely unassisted, and look honestly at the result. Physicians measure their unassisted detection rate. You should measure yours.
The people in Boxtown never had to be told any of this. When the stakes are your own air and your own electricity bill, effort recalibrates itself, and nobody in that room is confused about whether they are thinking.
The Missing Prime Directive
Automate the labor. Automate the drafting, the summarizing, the searching, the tedious reconciliation of forty tabs. I have no quarrel with any of it, and I would not give it back.
But keep one instrument in the house that the machine never calibrates. Some regular task, chosen by you, done cold and unaided, whose only purpose is to tell you the truth about where you actually stand. Everything else can be delegated. The measurement cannot, because a mind that has outsourced its own gauge has no way left to discover that it did.
Appendix: Goff’s Law, 2001
Since I have leaned on a twenty-five-year-old essay of my own and there is no longer a live copy of it anywhere on the internet, here it is. This ran in my Zen and Java series in 2001, on a personal site that predates this one by more than a decade. The Internet Archive never crawled the directory, so this reproduction is the only public copy. The generator tag in the original file reads Netscape 4.7 on SunOS 5.7, which tells you everything you need to know about the year.
The words are exactly as they ran. I have left the argument where it was and only the quotation marks have been modernized. My regular readers will recognize a phrase in the fourth paragraph that took me another two years to name properly.
Goff’s Law
From Zen and Java, copyright 2001, Max K. GoffPlease accept my humble apologies in advance for this; it is not my intent to advocate an elitist agenda or justification for social policies that do not honor the sanctity of the human spirit first and foremost above all other measures. I realize there is risk in discussing matters that might be viewed as politically incorrect. Please know that my intent is not to cause harm or injury to anyone; I consider myself to be egalitarian by nature and if anything I’m opposed to intelligence tests in the first place. There are many measures of intelligence in our species, only one type of which is the standardized IQ test, which reflects only one’s penchant for things cultural anyway. There is too much that is ineffable about our universe and the human condition to place much stock in measurements of native intelligence viz. the cultural hallucination we verbally share. Having said that….
Goff’s Law: By any measure of intelligence, half of humanity is and will always be below median IQ.
Since this observation is definitional it is presumably obvious and not meant to impugn or elevate any group or individual. It is simply true, based on statistical axiom. So how is this germane?
If there is a correlation between relative wealth or earning capability and IQ, which may be the case, especially in a knowledge-driven economy, then we must recognize that the “fitness landscape” will favor those on the high end of the median. Which is to say that unless we do something about it, the smart will inherit the earth. “Have vs. have not” trends will continue. At least that seems to be true on the surface. But in the Network Age, no one thing stands alone without the mediating influence of other factors, the unintended consequences, if you will.
According to the World Watch Institute, most of our planetary ecological malaise is due to the activities of the wealthiest 20% and the poorest 20% of the earth’s occupants. If we recognize the correlation between wealth and IQ, then we must also recognize the correlation between ecological malady and IQ. And since ecologies bind us all (as in, we’re all connected in the end anyway), the effects of activities that perturb the environment will be generalized to the world in toto. In other words, it doesn’t matter how smart you are or how much wealth you have if you can’t breathe the air, drink the water, break contaminated bread with your highly contagious neighbor, or enjoy the abundance of a planet forsaken. Nor does it matter whose fault it is. Clearly, the responsibility for remedy lies with each of us.
Most software developers are members of the other half. In fact, I dare say all are members of the other half. <humor> Their management may not be. </humor> But it is highly probable that if you are a software developer you sport an IQ greater than the median by any measure. That’s why you’re getting the big bucks. With that mantle, however, come responsibilities, not the least of which is the realization that the code we write must serve all, not just the wealthy and not just the smart. All. Half the adults on this blue bubble can’t even read. How can wealth propagate given these observations? Answering that question is key to understanding the scope of any effective vision we might share.
Leave a Reply