Think with AI. Keep the thinking.
Your memo is finished. The sentences are clean. The recommendation sounds decisive. Then someone asks why you rejected the alternative, and you reach for the chat history. A useful way to work with a large language model starts with that moment: what should you still be able to explain when the model is gone?
A finished page.
An unfinished thought.#
A document has an immediate result: someone can read it. Writing can also leave you with something less visible: a clearer model of the problem. Decide which of those outcomes your use of AI needs to protect.
In his Code to Care video, Don Woodlock describes a deliberate boundary: use LLMs to explore a subject, write and edit the first draft yourself, then bring the model back for grammar and spelling. This essay develops that approach into a practical workflow. It is a choice about how to work, not a scientifically established rule for every writer.[1]
Follow one example. You are preparing a pricing memo for Fieldnote, a fictional research assistant for small teams. You are considering a €19 plan per person, a €49 team plan, or metered usage. Those prices are illustrative proposals, not market data. Your task is to recommend a first experiment and explain what would make you change course.
An AI draft might recommend “a competitive, scalable pricing structure.” You still need to decide whose budget matters, how usage affects costs, and what evidence supports willingness to pay. Writing forces those unresolved choices onto the page—provided you actually make them.
AI can improve the page itself. In a randomized experiment involving 453 college-educated professionals, Noy and Zhang found that ChatGPT reduced time spent on selected professional writing tasks by 40% and increased rated output quality by 18%. That study measured task performance; it did not establish long-term learning effects.[2]
A high-school mathematics experiment by Bastani and colleagues found a gap between assisted practice and unaided performance. Compared with a control group, access to a general GPT-4 interface improved practice grades by 48%, but subsequent grades without AI were 17% lower. A tutor with learning safeguards raised practice grades by 127% and largely removed the later penalty, without demonstrating a positive unaided exam effect.[3]
Assisted practice grades
Later unaided grades
Assisted practice grades
These are relative differences in grades, not percentage-point changes or measurements of intelligence. The practical lesson is that the kind of help matters. Finishing practice successfully and acquiring a skill can come apart.
Essay writing: a limited signal
Kosmyna and colleagues’ preprint studied 54 participants in its first three sessions. LLM users reported less ownership and had more difficulty quoting their essays. The work does not establish permanent cognitive damage. A subsequent methodological critique raised concerns about sample size, EEG analysis, and reproducibility.[4][5]
Clinical AI: an association
An observational study at four Polish centres compared colonoscopies performed without AI before and after its introduction. Adenoma detection fell from 28.4% to 22.4%: six percentage points. This involved computer-aided polyp detection, not LLM writing. Its design cannot by itself establish that AI caused the decline.[6]
Bring that distinction back to Fieldnote. Your memo can look stronger while your understanding remains untested. You do not need to predict brain damage to make room for independent reasoning.
Make the question
bigger first.#
Give the model a research assignment with visible boundaries. Ask it to surface options, assumptions, and missing evidence. Keep the recommendation open until you have something worth comparing.
Start with your own rough position. For Fieldnote, write: “I lean toward a team plan because collaboration is the product’s main benefit. I do not yet know whether small teams share a budget.” This gives you an assumption to investigate before a fluent response supplies the framing.
Then ask for alternatives that change the economics: charge per person, per team, or per unit of usage. Ask who benefits and who objects under each. A longer list is useful only when it reveals a different mechanism or a neglected constraint.
Prompt · Open the decision
I am comparing pricing options for Fieldnote, a fictional
research assistant for small teams: €19 per person,
€49 per team, or usage-based pricing.
Do not recommend a winner yet.
For each option, identify the buyer, likely objection,
cost exposure, and assumption I would need to test.
Add one materially different alternative.
Separate supplied facts from your hypotheses.
If current sources are available, include the exact URL,
date, billing unit, and passage supporting each claim.
If you cannot verify something, mark it unverified.
Treat the response as a research map. Open the original sources yourself. A price without its billing period, usage allowance, region, or tax treatment is an incomplete comparison. A model without current browsing or supplied documents cannot verify today’s competitor prices.
A source states a plan’s price and billing unit. Save the page, date, and relevant passage.
You think a shared budget makes a team plan attractive. Record the reasoning separately.
You have no willingness-to-pay evidence for Fieldnote. Keep that gap visible.
Finish the exploration by asking for a serious objection: “Assume the team plan fails. What plausible mechanism explains the failure, and what observation would distinguish it from a weak product?” The answer might point toward high usage costs, purchasing friction, or low demand. Those are hypotheses to test, not facts produced by the prompt.
Set a stopping condition: you can explain the main options, their decisive trade-offs, and the next missing piece of evidence. Otherwise, asking for “more ideas” can become a way to postpone choosing.
Close the chat.
Make the case.#
The transition from options to a position deserves its own space. Keep your verified notes available. Put the generated prose aside. Write the recommendation in words that expose your reasoning.
Begin with three sentences: what you recommend, why, and what could make it wrong. For Fieldnote, a provisional position could be: “Test a capped €49 team plan. It matches our collaboration hypothesis and gives buyers a predictable bill. We have not established willingness to pay, so this is a pilot proposal.”
Now explain the rejected alternative. Per-person billing may collect more from large teams, but it could discourage inviting occasional collaborators. That trade-off is part of your argument. If you cannot explain it, you have found a question to resolve before polishing the wording.
The sentence you can defend
“We propose testing a capped team plan because shared use is central to the product.” You can identify the proposal, its rationale, and the assumption behind it.
The sentence that hides the work
“Our optimized pricing strategy will unlock sustainable growth.” Which price? Optimized against what? What evidence makes “will” appropriate?
A useful memo makes the chain explicit: observation, interpretation, recommendation, next test. It also preserves uncertainty. “We do not know the cost of heavy usage” is a reason to measure costs or cap a pilot. It is not a reason to ask for more confident prose.
Edit once for the argument
Read your draft without the chat beside it. Underline each factual claim and locate its source. Circle each recommendation and check its reasoning. Remove any sentence you cannot explain. Then read it aloud: where you need a long spoken qualification, the written claim may be too broad.
Self-review matters because trusting a tool and evaluating its output are different activities. A CHI study of 319 knowledge workers found that greater confidence in GenAI was associated with less reported critical-thinking effort. It was a survey, not a causal test of skill loss; the authors also describe work shifting toward verification and integration.[7]
This boundary should serve your purpose. If you need language or accessibility support, dictate your reasoning or work from your own outline. For routine production, a reviewed AI draft may be appropriate. When learning, original argument, or strategic judgment is the goal, reserve a meaningful part of the reasoning for yourself.
Clean the language.
Protect the claim.#
Bring the model back with a narrow brief and a visible record of changes. The useful output is a set of proposed edits you can judge, not a replacement document you must trust.
Suppose your draft says, “The team plan reduce purchasing friction, but the evidence is preliminary.” Changing “reduce” to “reduces” repairs subject–verb agreement. Removing the final clause changes how certain the claim appears. Turning it into “The team plan will accelerate adoption” introduces a different prediction.
Prompt · Make every change reviewable
Review this memo for grammar and spelling only.
Preserve its meaning, voice, numbers, citations,
qualifiers, and level of certainty.
Return a table: original phrase | proposed edit | reason.
Do not rewrite the whole document.
If a sentence is ambiguous, ask what I mean.
Flag any suspected factual problem separately;
do not silently correct or invent evidence.
The prompt sets a boundary; your review enforces it. Compare each change with the original. Pay special attention to may, will, some, all, and because. These small words carry scope, certainty, and causality.
That is a change of claim, not a spelling correction.
If you also want criticism of the argument, make it a separate pass: ask for unsupported claims and unanswered objections without requesting a rewrite. Decide which criticisms deserve a revision, then make that revision yourself. An editing brief should not silently become a decision brief.
Leave with more than the file
After finishing the Fieldnote memo, close it and write a short explanation of your recommendation, the strongest alternative, and the evidence that would change your mind. Check that explanation against the memo. Treat any gap as a prompt to revisit the reasoning, not as a diagnosis of your memory.
For your next important document, compare two outcomes: how efficiently you completed it and how well you can explain its choices. You can gain speed and retain ownership. The boundary is a deliberate allocation of work, not a contest to minimize AI use.
Who is doing the thinking?
Apply the method to the Fieldnote memo. Choose an answer, then inspect the reasoning.
1. The model gives you a competitor price with a link. What earns that claim a place in the memo?
Why this matters
B is the supported choice. Repetition does not supply independent evidence. Checking the source can reveal that a price is per person, billed annually, or unavailable under the conditions your memo assumes.
2. An edit changes “may reduce friction” to “will accelerate adoption.” How should you classify it?
Why this matters
A is the supported choice. The edit changes both certainty and the predicted outcome. Smoother language does not justify either change.
3. Your memo is convincing, but you cannot explain why you rejected per-person billing. What next?
Why this matters
C is the supported choice. The missing piece is an explanation. Use the verified notes to rebuild it; revise the recommendation if it no longer holds.
Follow the evidence.#
The studies below address different tasks and outcomes. They do not validate this exact three-stage workflow. Fieldnote, its proposed prices, prompts, and memo excerpts are teaching examples; no customer or competitor research is implied.
- How to Personally Use LLMs. Don Woodlock · Code to Care.Original video and transcriptThe starting point for the research–draft–edit division of labor. This article adds worked examples and independently checks the research claims.
- Experimental evidence on the productivity effects of generative artificial intelligence. Shakked Noy and Whitney Zhang · Science, 2023.Research articleRandomized professional writing experiment; time and rated output quality, not long-term learning.
- Generative AI without guardrails can harm learning: Evidence from high school mathematics. Hamsa Bastani and colleagues · PNAS, 2025.Research articleAssisted practice and unaided exams are distinct outcomes. The published correction concerns an author affiliation, not these findings.
- Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. Nataliya Kosmyna and colleagues · arXiv preprint, version 2, December 2025.PreprintSmall essay-writing study with behavioral and EEG measures. Used here with explicit limits on interpretation.
- Comment on: Your Brain on ChatGPT. Milos Stankovic and colleagues · arXiv preprint.Methodological critiqueRaises concerns about study design, reproducibility, and EEG methodology; it is not a replication study.
- Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Krzysztof Budzyń and colleagues · The Lancet Gastroenterology & Hepatology, 2025.Study abstractBefore–after comparison of procedures without AI. Computer-aided detection is a different technology and task from LLM-assisted writing.
- The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. Hao-Ping Lee and colleagues · CHI, 2025.Research publicationSelf-reports from 319 knowledge workers; associations and descriptions of work, not a causal estimate of cognitive decline.





