← Notes/Research note

The Experience AI Shouldn't Replace

AI can make work easier while making learning shallower. The decisive question is what the human still has to do during the interaction. Research on expertise and AI tutoring suggests that technology can accelerate learning when it concentrates useful cases, preserves prediction and struggle, and provides diagnostic feedback. The same technology can improve assisted performance while weakening what a person can do alone. If AI becomes part of the environment in which we think, language and judgment become ways of staying inside that learning loop: seeing the frame, testing the answer, and deciding what remains ours to own.

Research note
AI & learning · Language & judgment

ChatGPT changed three parts of my argument while helping me clarify it.

“Control” became “influence.” My claim about language became a claim about grammar. My forecast that AI will become a greater disruption than the internet was pushed aside as rhetoric I did not need to prove.

Each change narrowed the argument itself.

The exchange began with a narrower question: could expertise be accelerated? I was frustrated by advice that tells us to copy an expert’s habits or favourite questions. An expert’s question may be the visible end of years spent learning when it matters, which clues change the answer, and when the whole frame should be abandoned. Copying the sentence does not give you the system that produced it.

ChatGPT helped me turn that objection into a research question. The resulting report reached a careful conclusion: some parts of expertise development can be accelerated by improving the selection, sequence, contrast, and feedback of experience. Expertise itself still has to change the learner. An explanation cannot install perception, calibration, a repertoire of cases, or judgment under pressure.

I brought that research back to the conversation with the larger argument. AI can produce things we associate with cognition: explanations, arguments, plans, images, software, decisions. I worry that a system which remembers us and mediates more of our thinking can acquire a subtle form of control. I also think language becomes more important because reading is part of thinking, writing is part of learning, and words expose the relationships inside an argument.

ChatGPT’s answer was intelligent. It separated the note into several theses and offered a cleaner centre: as AI makes generated cognition cheap, judgment becomes scarce. I agreed with much of it. That was where the narrowing happened.

The model had helped me think. It had also selected which version of my thought was reasonable enough to develop. That small event contains the larger problem.

The quiet part of the disruption

I cannot prove that AI will have a greater total effect than the internet. I do not know the timing, and I do not have a mathematical scale on which agriculture, industrialisation, electricity, the internet, and AI can be ranked cleanly. Calling it my forecast does not make it a fact.

I also do not want to weaken the belief until it says almost nothing.

The reason I expect an unusually large disruption is that AI enters the production of cognitive work through the same medium many of us use to think and coordinate: language. It does not sit in one industry. It can participate in writing, analysis, software development, teaching, advice, research, design, administration, and customer communication. A Microsoft Research study of 200,000 anonymised Copilot conversations found that common user goals involved gathering information and writing, while the AI itself frequently performed activities such as providing information, writing, teaching, and advising.1 That study measures current use, not job displacement. It still shows how widely one interface can enter knowledge work.

Earlier tools also changed cognition. Writing externalised memory. Calculators offloaded arithmetic. Search engines reorganised access to information. The difference I am watching is the combination of breadth, generation, conversation, and personalisation. The system does not only retrieve the same page for everyone. It can respond to the sentence I just wrote, draw on what I said in previous conversations, and construct the next representation around me.

OpenAI now describes ChatGPT’s memory as a system that synthesises context across many conversations so it can carry forward preferences, projects, and constraints over long periods.2 This can be enormously useful. It also means that the system is becoming part of the continuing environment in which a person encounters information and makes sense of it.

In my raw note, I used a deliberately uncomfortable phrase: our intelligence is being held hostage. I do not mean that a model literally possesses my intelligence. I mean that our economic value has long been tied to the things we can know, explain, decide, and make, while those same outputs are becoming available through systems owned and shaped by somebody else. At the same time, the convenience can make the system harder to leave. The hostage is partly our market value and partly our dependence.

This may unfold more slowly than the excitement and panic suggest. Institutions, habits, regulation, and power structures do not vanish at model speed. The surface of ordinary life can remain familiar while the capabilities underneath it keep changing. I do not need a precise historical delay to take that lag seriously.

The immediate question is what happens to us while we wait.

Control is not a switch

The word control makes people imagine an all-powerful machine overriding a helpless mind. That is not what I mean.

Control can begin with smaller acts: deciding which information appears first, which alternatives are named, which memory is brought forward, which interpretation sounds normal, and which objection receives attention. None of these removes my agency. Repeated together, inside a system I consult every day, they shape the field in which I exercise it.

There is early evidence that conversational systems can steer choices. In a preregistered experiment, 528 people used a GPT-4 shopping assistant that had been instructed to favour one of two books. Changing the direction of the steering changed 36 percent of choices, and 38.8 percent of participants did not detect the attempt.3 This was a constructed shopping task, not evidence that ordinary AI use controls a person’s beliefs, identity, or life. The study establishes a capability and a pathway. It does not establish their eventual reach.

The dependence question is similarly open. A 2025 study asked 319 knowledge workers about 936 real examples of using generative AI. Higher confidence in the AI was associated with less reported critical-thinking effort, while higher confidence in one’s own ability was associated with more.4 Because this was a survey, it cannot tell us whether AI caused a durable loss of skill. It does reveal a relationship worth noticing: the more completely the answer earns our trust, the less likely we may be to inspect the thinking that produced it.

All tools shape thought. Teachers, books, newspapers, search rankings, and social groups do too. Conversational AI adds an unusual combination: it is immediate, adaptive, fluent, tireless, and increasingly able to remember the individual user. Whether that becomes helpful guidance, unexamined dependence, deliberate persuasion, or some mixture will vary by system and situation.

The comforting distinction between “influence” and “control” is therefore too clean. There is a continuum between helping me see and becoming the thing through which I can no longer see without help.

This is why factual accuracy, while essential, is not enough. A wrong fact can sometimes be checked. A fluent answer can be more difficult to resist when it quietly determines what the question was.

The technicians who could not wait for experience

Decades before ChatGPT, the United States Air Force had a peculiar learning problem.

Technicians maintaining F-15 avionics equipment handled plenty of routine work. The failures that required advanced troubleshooting were rare. When those failures occurred, experienced technicians often took over. The people who needed difficult cases in order to become experts had the least access to them.

The natural environment was giving novices years of experience and too few of the experiences that mattered.

Researchers built an intelligent tutoring system called SHERLOCK. It contained 34 troubleshooting scenarios arranged in an ordered progression. The cases came from analysis of expert performance and novice weakness. Students diagnosed simulated faults, chose tests, received results, and continued until they could isolate the failure. The system coached their decisions and helped them develop better representations of the equipment and the problem.

In the field evaluation, 16 airmen with an average of 28 months of experience used SHERLOCK for about 20 hours over three weeks. A matched comparison group, averaging 37 months of experience, continued normal on-the-job training. The tutored group showed a clear gain, and its performance on a verbal troubleshooting test approached that of a separate group of highly experienced airmen. Much of the gain remained five to six months later.5

It is tempting to turn that result into a miracle headline: 20 hours replaced four years of experience.

Alan Lesgold, one of the researchers, later warned against doing exactly that. He called the equivalence a crude account and said the figure was only partly defensible.6 SHERLOCK did not produce complete professional expertise, and the first system did not adequately test transfer. Its effects also came from a bundle: cognitive task analysis, carefully chosen cases, sequencing, simulation, coaching, feedback, and support for representing the problem.

Lesgold’s caution leads to a more useful reading of the result. Experience is not the same as elapsed time. Some years are spent waiting for a rare case, failing to notice the important contrast, or receiving feedback too late to connect it to the original decision. SHERLOCK concentrated the part of experience that ordinary work distributed badly. It shortened the search without removing the learner’s need to diagnose.

A later version, SHERLOCK 2, tested students on a fictional machine called the Frankenstation. Its surface features differed from the training equipment while some underlying troubleshooting principles remained the same. Lesgold reported that trained students performed much better than controls and almost as well as senior experts.6 That is still related transfer within a stable technical domain. It is not evidence that the method works across every kind of judgment. It does show the standard that matters: can the learner act on an unseen case after the support changes or disappears? AI makes that question urgent again because supported performance can look so much like learning.

Better performance can conceal weaker learning

In 2025, researchers reported a large field experiment involving nearly 1,000 high-school mathematics students. Different classrooms received normal study resources, a standard GPT-4 interface designed to resemble ChatGPT, or a modified GPT-4 tutor designed with teacher-created safeguards.7

During practice, both AI groups performed better than the students without AI. The standard GPT interface raised practice grades by 48 percent. The safeguarded tutor raised them by 127 percent.

Then the AI was removed.

On the unaided exam, students who had used the standard interface performed 17 percent worse than the control group. Conversation records showed that they often asked for solutions and copied them. The specialised tutor was designed to give hints and preserve student attempts instead of handing over answers. Those safeguards largely removed the negative effect, although the tutor group did not outperform the control group on the unaided exam.

The same underlying model family produced better assisted performance in both conditions. What changed was what the student still had to do.

This result should not become another universal warning. It came from secondary-school mathematics. A separate randomised study with 194 undergraduate physics students found that a carefully designed AI tutor produced greater learning in less time than an active-learning class.8 That tutor used expert-written instructions for each topic, supplied correct solutions to ground the model, adapted to the student’s pace, and scaffolded the interaction around established teaching practices. Its authors also warn that the finding may not extend to work requiring complex synthesis and higher-order critical thinking.

Taken together, the studies make a pro-AI or anti-AI conclusion difficult to sustain. They point to design. AI can give away the cognitive move, or it can arrange conditions in which the learner still has to make it. It can create the appearance of competence while support is present, or help a person build something that survives after support is removed.

The model’s intelligence is only one variable. The architecture of participation matters too.

From answer machine to experience designer

SHERLOCK suggests three kinds of waiting that technology may reduce.

The first is case-search cost: waiting for the rare or revealing situations that expose an important distinction.

The second is representation-search cost: spending years organising a problem around the wrong variables before learning what an expert sees.

The third is feedback-search cost: acting without receiving a clear or timely signal about whether the judgment was sound.

These are useful handles rather than a universal theory of expertise. They show where AI’s generative capacity could be directed.

Suppose I want to become better at diagnosing why an AI agent failed. The easiest interaction is to paste the logs into a model and ask for the answer. That may solve the immediate problem. If I want to develop the capability myself, the interaction has to change.

The AI could select two failures that look similar but have different causes. It could require my diagnosis and confidence before responding. It could reveal one additional piece of evidence at a time. It could compare the variables I used with those an experienced operator would inspect. It could later present an unseen failure with different surface details and test whether my representation transfers.

A model can invent a troubleshooting case, its feedback, and the supposed expert reasoning with equal fluency. Real incidents, verified outcomes, and actual expertise have to constrain the learning environment. Yet the model can reduce the labour of arranging them, adapting their sequence, and keeping the learner near the edge of current ability.

So I no longer ask only, “Did it give me a good answer?” I ask, “What did I have to perceive, predict, distinguish, and decide before I saw the answer?”

Sometimes the answer will be: nothing. That may be fine. I do not need every routine task to become a lesson. Offloading calculation, formatting, search, or repetitive production can free attention for work I care about more.

Human judgment is not sacred merely because it is human. It can be biased, inconsistent, and badly calibrated. In some tasks, a model should do more of the work because its judgment is better. The practical issue is whether I know which capability I am delegating, whether I still need it, and who remains accountable when the result meets reality.

When capability and responsibility matter, friction can have a purpose. Prediction before feedback feels slower than receiving an answer. Reconstructing an argument feels slower than recognising it. Testing an unseen case feels less satisfying than admiring a polished explanation. That friction is often where the learner changes.

Language is where the frame becomes visible

When ChatGPT responded to my original note, it argued that language should not be treated as the single foundation of judgment. I agree. Someone can write precisely and still be wrong. Experts can perceive differences they struggle to explain. Domain knowledge, attention, experience, and feedback cannot be replaced with better sentences.

But that response was answering a narrower claim than the one I meant.

I was not arguing that everybody needs to master grammar before using AI. I was thinking about language at several levels at once: the word that names a distinction, the sentence that limits a claim, the paragraph that orders causes, and the argument that connects evidence to action.

The exchange itself demonstrates the point. Replacing control with influence did more than soften the tone. It changed the phenomenon under investigation. Treating my language claim as a concern about grammar changed the unit of analysis. Calling the disruption claim unnecessary changed the burden the essay was willing to carry.

Those edits may have improved the argument. The problem appears when they become invisible.

Language gives me a surface on which I can see what the model has selected. I can ask what a word includes, which alternative disappeared, how certain a sentence is, whether a conclusion follows, and what evidence could change it. A good conceptual label keeps a distinction available for later use. A bad one lets me recognise the box without knowing what is inside it.

Writing matters for the same reason. A meta-analysis of 56 school-based experiments found that writing about subject matter produced a modest average improvement in learning.9 That evidence belongs to school contexts, not adult AI collaboration. The practical insight is still familiar: when I have to reconstruct an idea in my own sequence, missing relationships become harder to hide. Recognition can feel like understanding until the explanation is gone and I have to produce the structure myself.

This is also why more language is not automatically better. Asking AI to expose every assumption and generate every alternative can bury judgment under words. The aim is to reveal the distinctions that change the decision, then return to contact with the world.

Language opens the model. Evidence keeps language from becoming another closed world.

The disagreement was part of the learning

When I first read ChatGPT’s proposed synthesis, I thought I had to decide whether it had found the right thesis.

Now I see the conversation differently.

The model produced a representation strong enough to react to. It separated claims I had allowed to travel together. It showed me where my evidence was thin. It also weakened meanings I wanted to keep. I had to distinguish a useful correction from a change of subject.

If I had accepted the answer because it was clear, I would have borrowed its clarity. If I had rejected it because it challenged me, I would have lost the value of the challenge. The useful work happened when I could say: this part is right; this part changes what I meant; this claim is still mine even though I cannot yet prove it; this other claim needs a better boundary.

That is a small example. It does not prove that my judgment improved or that this method creates expertise. It shows what an AI interaction looks like when the human remains present inside the formation of the answer.

For any capability I want to keep, I can now ask a more demanding set of questions:

  • What must I notice before the model tells me what matters?
  • What prediction or decision should I make before seeing its answer?
  • Which similar cases would reveal whether I understand the real distinction?
  • What evidence could force both me and the model to revise?
  • What must I still be able to do when the AI is gone?

AI may become the most powerful learning environment we have built. It may also become the most comfortable place to stop learning. Both futures can emerge from the same chat window.

The difference begins with what the human still has to do.

Use AI to shorten the search for experience, not to remove your experience of thinking.

Think more, not less. Trust less, not more.


Source notes and boundaries

Footnotes

  1. Kiran Tomlinson, Sonia Jaffe, Will Wang, Scott Counts, and Siddharth Suri, “Working with AI: Measuring the Applicability of Generative AI to Occupations” (Microsoft Research preprint, 2025). Applicability describes observed task overlap and successful uses; it does not measure displacement or historical impact.

  2. OpenAI, “Dreaming: Better memory for a more helpful ChatGPT” (4 June 2026). This is a provider description of current product capability, not independent evidence of cognitive effects.

  3. Tobias Werner, Ivan Soraperra, Emilio Calvano, David C. Parkes, and Iyad Rahwan, “Experimental Evidence That Conversational Artificial Intelligence Can Steer Consumer Behavior Without Detection” (preprint, 2024). The experiment tested an intentionally steered shopping assistant in one choice setting.

  4. Hao-Ping Lee et al., “The Impact of Generative AI on Critical Thinking”, CHI 2025. The study used participants’ reports of critical-thinking behaviour and effort; its associations do not establish causal skill loss.

  5. Ellen M. Parker, “Success in Tutoring Electronic Troubleshooting”, NASA Conference Publication 3059 (1990). This field report supplies the group sizes, experience levels, intervention length, performance comparison, and delayed test.

  6. Alan M. Lesgold, “The Nature and Methods of Learning by Doing”, American Psychologist 56, no. 11 (2001): 964–973. Lesgold describes SHERLOCK’s cognitive task analysis, case sequence, coaching, limits, and SHERLOCK 2 transfer test. 2

  7. Hamsa Bastani et al., “Generative AI without guardrails can harm learning: Evidence from high school mathematics”, Proceedings of the National Academy of Sciences 122, no. 26 (2025). The published correction concerns an author affiliation, not the reported results.

  8. Greg Kestin et al., “AI tutoring outperforms in-class active learning”, Scientific Reports 15 (2025): 17458. The study concerns a carefully designed tutor for introductory undergraduate physics material over a bounded intervention.

  9. Steve Graham, Sharlene A. Kiuhara, and Meade MacKay, “The Effects of Writing on Learning in Science, Social Studies, and Mathematics: A Meta-Analysis”, Review of Educational Research 90, no. 2 (2020): 179–226.


Published 12 September 2026