by admin | Aug 27, 2026 | Research & Productivity
A few months ago I caught myself doing something odd. I had used AI to draft a page for one of my sites, published it, and a week later I could not have explained the argument in it without opening the file. The work was done. The thinking, apparently, had not stuck.
That feeling has a name now, and researchers have started measuring it. The question of AI and critical thinking is no longer just a worry people raise at dinner. Three studies, two of them peer reviewed, have looked at what happens inside the heads of people who lean on AI, and the results are worth knowing before you build your next work habit around a chatbot.
The short version: the risk is real, it is not automatic, and the difference comes down to how you use the tool rather than whether you use it.
What the research says about AI and critical thinking
The most useful study of working adults came from Microsoft Research and Carnegie Mellon University, presented at CHI 2025. The team surveyed 319 knowledge workers and collected 936 first-hand examples of generative AI used in real work tasks.
The headline finding is a confidence effect, and it cuts both ways. People with higher confidence in the AI did less critical thinking. People with higher confidence in their own ability at the task did more. Same tool, opposite outcomes, depending on who was holding it.
The researchers also found that the thinking does not disappear so much as move. It shifts away from producing an answer and toward verifying information, stitching AI output together, and managing the task overall. That is still real cognitive work. It is just different work, and it only happens if you choose to do it. You can read the full paper on the Microsoft Research site.
The MIT essay study, and what it does not prove
The study that got the loudest headlines came from the MIT Media Lab. Researchers put EEG caps on participants and asked them to write essays under three conditions: using an LLM, using a search engine, or using nothing at all.
Brain connectivity was strongest in the no-tools group, moderate in the search group, and weakest in the LLM group. The LLM writers also reported the lowest sense of ownership over their essays, and struggled to accurately quote work they had produced minutes earlier. The authors called the pattern “cognitive debt”. The full paper is on arXiv.
Now the honest part, because this study has been badly oversold online. It involved 54 participants across the first three sessions and only 18 in the fourth, and it looked at one narrow task. It is still a preprint. The authors revised it in December 2025, but it has not been through peer review. That is a signal worth taking seriously, not a verdict on your brain.
The order you work in seems to matter
Buried in the same MIT study is the most practical finding of the lot, and almost nobody reported it.
In the final session the researchers swapped the groups around. People who had been writing unaided and were then given an LLM showed higher memory recall and stronger brain activation, closer to the search-engine group. People who had been leaning on the LLM and then had it taken away showed reduced connectivity and looked under-engaged.
Read that again, because it is the whole practical lesson. Thinking first and bringing AI in second looked healthy. Starting with AI and trying to think afterwards did not.
The 2026 study that names the difference
The newest piece of this puzzle is also the most useful, and it arrived in July 2026 in the peer-reviewed journal Frontiers in Psychology. Researchers surveyed 589 university students and early-career knowledge workers across three rounds spaced out over time, and they split AI use into two modes instead of treating it as one thing.
Dependent offloading is handing over the actual thinking: accepting the output with little examination, letting the AI structure your ideas, treating what it produces as the finished article. Autonomous offloading is using the same tool as a scaffold: taking the output as a starting point, comparing it against your own reasoning, and keeping ownership of the final result.
The two modes pulled in opposite directions. Dependent use was linked to handing over cognitive agency and to lower motivation, and through those to weaker self-rated judgement, creativity and depth of thinking. Autonomous use was linked to higher motivation and better self-rated outcomes. Same tools, same tasks, opposite results based only on how people engaged.
Then comes the finding that should change how you work. Both modes produced comparable immediate benefits. The session feels equally productive either way, which means unhelpful AI use is very hard to spot from experience alone. Nothing in the moment tells you which mode you are in. You can read the full study on the Frontiers in Psychology site.
Keep it in perspective. It asks people to rate their own thinking rather than testing it, so it shows associations, not proof of cause. But it is peer reviewed, far larger than the MIT study, and it names the distinction the other two keep circling.
Five habits that protect your thinking
- Write your rough version first. Even five messy bullet points before you open the chat window changes the whole session. You arrive with a position instead of asking to be given one.
- Ask the AI to question you rather than answer you. Try “argue against this” or “what am I missing here” instead of “write this for me”. You get pushback rather than a finished product you never examined.
- Verify anything you would be embarrassed to get wrong. Names, numbers, dates, citations, legal or medical claims. AI systems produce confident wrong answers regularly, which is why our guide to AI hallucinations is one of the most useful things to read before you trust an output.
- Explain the result out loud in your own words. If you cannot, you have not understood it, and you will not be able to defend it in a meeting or a viva.
- Notice when a task feels low-stakes. The Microsoft researchers found that people review AI output far less carefully when they judge a task to be unimportant. Low-stakes tasks are exactly where sloppy errors slip through into public work.
Important tip: if you cannot explain the answer to another person in your own words, you have not learned it. You have only borrowed it.
When letting AI do the work is completely fine
None of this means you should feel guilty every time you open a chatbot. Offloading is only a problem when you are offloading something you actually needed to learn.
Reformatting a messy list, fixing spelling, converting a table, writing boilerplate you have written a hundred times, summarising a document you are only skimming for one fact. Hand all of that over without a second thought. From my own experience running websites and online tools, that category is where AI saves genuine hours, and none of those hours were making me smarter.
The line is simple. If the task is the learning, do it yourself first. If the task is friction around the learning, automate it. That distinction is also why our post on using AI tools without cheating keeps coming back to the same test.
What this means for students and researchers
If you are studying, the stakes are higher, because the entire point of the work is to build something in your head that stays there. A summary you did not read leaves nothing behind. Our plan for using AI to study for exams is built around this: generate questions, not answers.
Microsoft Research has since argued for designing AI as a “tool for thought” rather than an assistant that hands you finished work, warning that people risk becoming validators of machine output rather than authors of their own. You can read their argument here. Until the tools are built that way, the burden of using them well sits with you.
Common Questions
Does using AI actually make you less intelligent?
No study has shown that. What the research shows is reduced mental effort during AI-assisted tasks and weaker recall of AI-assisted work. Those are measurable short-term effects on specific tasks, not evidence of permanent change.
Is it better to just avoid AI for serious work?
Avoiding it entirely costs you speed and puts you behind on a skill employers now ask for. The evidence points toward using it deliberately rather than avoiding it, which is also the theme of our guide to the AI skills that matter most for future jobs.
What is cognitive offloading?
It means using something outside your head to do mental work for you. A shopping list, a calculator and a satnav all count. AI is a far broader version of the same thing, which is why where you draw the line matters more.
Final takeaway
AI and critical thinking are not enemies. The research suggests they come apart only when you hand over the part of the task that was supposed to change you. Think first, use the tool second, verify what matters, and be able to explain the result without the screen in front of you. Do that and the tool stays a tool.
by admin | Aug 6, 2026 | Research & Productivity
Exam week has a familiar shape. Twelve lectures to revise, four days left, and a folder of notes you have not opened since term started. So you paste a topic into a chatbot, read a clean summary, and feel like you have done something. Then you sit the paper and realise you recognised the material without being able to produce it.
Recognising an answer and being able to write one are two different skills. Exams only test the second. The good news is that if you use AI to study for exams properly, the newer study features from OpenAI and Google are built for exactly that gap.
What it really means to use AI to study for exams
There are two ways to use AI to study for exams. The first is as an answer machine: you ask, it explains, you read, you move on. It feels fast and teaches you very little.
The second is as a tutor that makes you do the work. It asks you questions, waits for your answer, corrects you, and only then moves forward. That version is slower and far more useful. Both OpenAI and Google now ship a mode built around this idea, and both describe it the same way: guide the student rather than hand over the solution.
From my own experience building websites and learning technical tools, the pattern holds outside exams too. Reading documentation feels productive. Trying to explain the thing to somebody else is when you find out what you actually understood.
Step 1: turn your syllabus into a plan before you ask for answers
Before you revise anything, give the AI your reality. Paste in the topic list or the module outline and tell it how many days you have, how many hours a day you can realistically give it, and which topics scare you most.
Ask for a day by day plan that spends more time on your weak topics and puts a short review of older topics at the start of each day. Then edit it. A plan you did not adjust is a plan you will abandon by Tuesday.
If time management is your bigger problem, our guide on how to use AI for time management and daily planning covers the wider workflow.
Step 2: switch on a study mode instead of using plain chat
OpenAI launched study mode in ChatGPT in July 2025. You turn it on by selecting Study and learn from the tools menu, then asking your question. Instead of answering outright, it uses guiding questions, hints, scaffolded explanations and short knowledge checks, and you can toggle it off mid conversation if you just want the answer. OpenAI is open about the trade off: it runs on custom instructions, so it can behave inconsistently or make mistakes.
Google has an equivalent called Guided Learning in the Gemini app. It asks you questions back, breaks a problem into steps, adapts to your level, and can build a study guide from course material you upload. Gemini also generates flashcards and quizzes, and Google says the mode runs on LearnLM, a version of its models tuned for learning.
The practical difference is real. Plain chat gives you a paragraph to read. Study mode gives you a question to answer, which is the thing your brain will need to do in the exam hall.
Step 3: make flashcards and quizzes from your own notes
Generic flashcards off the internet cover a generic syllabus. Yours does not. This is where NotebookLM earns its place: you upload your lecture notes, slides or readings, and it generates flashcards and quizzes grounded only in those sources. You can set the topic and the difficulty, share a set with classmates by link, and click explain on any card to get a fuller answer with citations pointing back to your original document.
That citation link matters more than it sounds. When a flashcard looks wrong, you can check it against your own slide in one click instead of guessing. We wrote a full walkthrough of how to use NotebookLM if you have not tried it yet.
NotebookLM also has audio formats now, including a short Brief summary and a Debate format where two AI hosts argue different sides of a topic. That one is useful for essay subjects where you need to hold two positions in your head.
Step 4: explain it back before you move on
After each topic, close the notes and explain it to the AI in your own words, out loud or typed. Then ask it to point out what you left out or got wrong.
This is the cheapest high value habit in the whole list. It takes three minutes, it needs no special tool, and it is brutally honest. If you cannot explain photosynthesis or a discounted cash flow without looking, you have not learned it yet, no matter how many summaries you read.
Step 5: practise under something like exam conditions
Give the AI a real past paper question or ask it to write one in your exam format, then answer it with a timer running and nothing open. Afterwards, paste your answer back and ask for marking against the actual marking criteria if your course publishes them.
Ask for the two specific things that would raise the grade rather than a general comment. Vague feedback is easy for a model to produce and useless to you.
A five day plan you can copy
- Day 1: Build the plan, upload your notes, generate flashcards for the two weakest topics.
- Day 2: Study mode on your weakest topic, then explain it back with the notes closed.
- Day 3: Second weakest topic, plus a ten minute flashcard review of day 2.
- Day 4: One timed past paper question, marked and reviewed. Fix the gaps it exposes.
- Day 5: Quiz yourself across everything, review only what you get wrong, then stop early and sleep.
Where AI still gets things wrong
AI models state wrong things confidently, and a flashcard is a very confident format. If a card contradicts your lecturer, your lecturer sets the exam. Trust the source, not the summary. Our post on why AI sometimes gives wrong answers explains why this happens.
Important tip: only generate study material from sources you have actually uploaded, and check anything you plan to memorise against your own notes at least once. A wrong fact you drilled twenty times is worse than a gap you knew about.
Two other cautions worth a minute of your time. Be careful what you upload, especially unpublished material or anything belonging to somebody else, and check your institution rules before you paste coursework anywhere. Working through years of websites and online tools has taught me that the upload button is the easiest place to make a quiet mistake. Our guide on using AI tools without cheating covers the academic integrity side properly.
Common Questions
Is using AI to study for exams cheating?
Using AI to quiz yourself, plan revision or explain a concept is studying, not cheating. Submitting AI written work as your own is a different thing entirely. When in doubt, read your institution academic integrity policy, because the rules vary between universities and even between modules.
Do I need to pay for these study features?
Not to get started. OpenAI made study mode available to logged in users on its free tier as well as paid ones, Guided Learning is in the Gemini app, and NotebookLM has a free tier. Paid plans mainly raise usage limits, so try the free versions before you spend anything.
Can AI predict what will be on my exam?
No, and be suspicious of anything that claims otherwise. It can spot themes across past papers you give it and generate practice questions in the same style, which is genuinely useful, but it has no knowledge of your unseen paper.
Final takeaway
The tools have quietly got better at the one thing that matters here. Study mode, Guided Learning and NotebookLM flashcards all push you to retrieve the answer rather than read it. That is the whole trick.
Pick one topic today, put the AI in study mode, and let it question you for fifteen minutes. If you finish that session slightly uncomfortable, it is working. For a wider view of how these tools fit into study and work, start with our guide on how AI can help with research and productivity.
Useful official sources: OpenAI on study mode, Google on Guided Learning in Gemini, and Google on NotebookLM flashcards and quizzes.
by admin | Jul 30, 2026 | Research & Productivity
You export a report, open it, and four thousand rows stare back at you. The answer your manager or your supervisor wants is somewhere in that file. Finding it used to mean an afternoon of pivot tables and half remembered formulas.
This is one area where AI has become properly useful. You can hand a spreadsheet to a chatbot, ask plain questions about it, and get back a summary, a chart, and the formula you were trying to remember. Below is how to analyze data with AI using free tools, which questions actually work, and where these tools still get things wrong.
What it means to analyze data with AI
Two different things share the same name, and mixing them up causes a lot of confusion.
- Chat tools that read a file you upload. You give ChatGPT, Claude, or Gemini a CSV or Excel file and ask questions about it in normal language. This route is free to try.
- AI built into the spreadsheet itself. Copilot in Excel and Gemini in Google Sheets sit inside the app and can edit your sheet directly. Both normally need a paid plan.
If you are starting out, begin with the first one. It costs nothing and it teaches you what these tools are good at before you pay for anything.
The free way: upload your spreadsheet and ask
The workflow is short. Tidy the file, upload it, then ask a specific question.
- Tidy the sheet first. One header row, one record per row, clear column names, no merged cells or decorative blank rows. OpenAI says the same thing in its own guidance: structured data with clear column names and one record per row gives the best results.
- Save it as CSV or XLSX. Delete the columns you do not need before you upload.
- Upload and say what you want to learn. Not “analyze this”, but “which three products lost the most sales between March and June, and by how much”.
If you have already tried our guide on how to chat with a PDF using AI, this will feel familiar. Same idea, different file type.
Which free tools handle spreadsheets
ChatGPT can inspect an uploaded file, answer questions about it, and build tables and charts. Free accounts are limited to three file uploads per day, and a CSV or spreadsheet can be roughly 50MB at most, as set out in the official data analysis help article.
Claude takes a slightly different approach. Its analysis tool writes and runs real code to do the calculation instead of predicting an answer as text. For anything involving actual arithmetic, that distinction matters more than it sounds.
Gemini works the same way in the main Gemini app, and it connects neatly to Google Drive if your files already live there.
The built-in option inside Excel and Google Sheets
If your workplace or university already pays for one of these, the AI is sitting in the app and most people never open it.
Copilot in Excel lives on the Home tab of the ribbon. Microsoft lists four main jobs for it: importing data, highlighting and sorting and filtering, generating formulas and explaining how they work, and surfacing insights as charts, PivotTables, summaries, trends, or outliers. It needs an eligible Microsoft 365 subscription, and your data has to be formatted as a table or supported range before Copilot can read it.
Gemini in Google Sheets opens from the Ask Gemini button in the top right. It can create tables and formulas, generate analysis and charts, and carry out edits you describe in words, including conditional formatting, pivot tables, sorting, filtering, and multi step cleanups. Two practical notes from Google: it needs an eligible Workspace or Google AI plan, and it works best on native Sheets files, so an uploaded .xlsx should be converted with File and then Save as Google Sheets. If a cell shows a formula error, hovering over it and clicking Fix asks Gemini to explain what went wrong.
Five questions worth asking about any dataset
Vague prompts get vague summaries. These five are a reliable starting set, and they work in every tool above.
- Describe this dataset. How many rows are there, what does each column mean, and what data is missing?
- What are the top five and bottom five rows by [column], and what is the gap between them?
- Group this by month and tell me whether the trend is rising or falling.
- Which rows look like duplicates, typing errors, or outliers?
- Write the Excel formula that does this, and explain each part of it.
Important tip: ask for the working, not just the answer. Adding “show me the steps and the columns you used” turns a black box into something you can actually check, and it is the single habit that catches most mistakes.
If your prompts keep returning generic replies, our guide on how to write better AI prompts covers the structure in more detail.
Always check the numbers before you use them
This is the part people skip. A chatbot answering in plain text can produce a confident figure that is simply wrong, which is the same problem behind AI hallucinations. Tools that run code are safer, but code can still read the wrong column or quietly ignore blank rows.
Three checks take about a minute between them. Recalculate one figure by hand. Confirm the row count the AI reports matches the row count in your file. Ask which columns it used, and read that answer properly.
From my own work on websites and analytics data, the failure is almost never the arithmetic. It is the tool confidently answering a slightly different question than the one I asked, and it looks right until someone checks.
Keep private data out of the chat
Spreadsheets are usually the most sensitive files an organisation owns. Customer lists, student records, salaries, and health data should not be dropped into a personal chatbot account just because it is quicker.
A safer habit: delete the name, email, and ID columns before you upload, or build a small anonymised sample that keeps the shape of the data without the identities. Our guide on how to use AI safely and protect your privacy walks through the account settings that matter here.
Common Questions
Do I still need to learn Excel? Yes, and probably more than before. AI writes the formula, but you are the one who has to notice when the result looks wrong. Understanding what a VLOOKUP or a pivot table does is what makes you useful next to the tool rather than dependent on it.
Which free tool is best for spreadsheets? For calculation heavy work, a tool that runs code is the safer bet. For quick summaries and charts, any of the three will do. Try the same file in two of them and compare the answers, which is also a decent accuracy test.
Can AI make charts from my data? Yes. All the tools mentioned can produce charts, and both Copilot in Excel and Gemini in Sheets can insert one directly into the sheet. Google notes that a chart Gemini inserts sits on a new tab with its own copy of the data, so it does not update when the original figures change.
Final takeaway
You do not need a data science course to get value here. Tidy one spreadsheet, upload it to a free tool, ask it the five questions above, then verify one number by hand. That single round trip will teach you more about what AI can and cannot do with data than any amount of reading.
Start small, keep the sensitive columns out, and treat every figure it gives you as a draft until you have checked it.
by admin | Jul 23, 2026 | Research & Productivity
Ask an AI chatbot a question and you get an answer in seconds. That speed is great for quick facts, but it falls apart the moment you need real depth. A seminar paper, a market overview, a thesis chapter. For that kind of work, a single fast answer is never enough.
This is the problem deep research modes were built to solve. Instead of replying instantly, the AI goes away for a few minutes, reads dozens or even hundreds of web sources, and comes back with a long, structured, cited report. In this guide I will explain what AI deep research actually is, how the main tools compare, and how to use them without getting burned.
What Is AI Deep Research?
AI deep research is an agent mode inside tools like ChatGPT, Gemini, and Perplexity. You give it one detailed question, and instead of answering from memory it plans a research strategy, runs many web searches, reads the sources it finds, and writes a report with citations you can check.
The difference from normal chat is time and effort. A regular answer takes seconds and often comes from the model’s training data. A deep research run can take anywhere from a few minutes to half an hour, because the AI is actually browsing and reading before it writes.
How It Works, Step by Step
- Plan: the AI turns your question into a research plan. Gemini even shows you this plan first so you can edit it before anything runs.
- Search and read: it runs many searches and opens the pages, the way you would with thirty browser tabs, just much faster.
- Reason: it compares sources, notices gaps, and searches again to fill them.
- Report: you get a structured document with sections and citations, ready to verify and reuse.
The Three Main Tools Compared
ChatGPT deep research is the heavyweight. OpenAI describes it as an agent that finds, analyzes, and synthesizes hundreds of online sources into a report at the level of a research analyst. Runs can take tens of minutes, and the reports are usually the longest and most detailed of the three.
Gemini Deep Research stands out for control. It shows you a multi-point research plan before it starts, can browse hundreds of websites, and can even turn the finished report into an audio overview you can listen to on a walk.
Perplexity Deep Research is the fast one. It typically finishes in two to four minutes, performing dozens of searches and reading hundreds of sources. It is also available on the free plan with a limited number of runs per day, which makes it the easiest way to try this kind of tool.
What It Is Good At, and Where It Fails
Deep research shines at mapping a topic you are new to: finding the main sources, the key debates, and the vocabulary of a field. It is excellent for background sections, tool comparisons, and market or policy overviews.
It is not a replacement for reading. The reports can still contain errors, and citations always need checking before anything goes into your own work. I covered this problem in detail in my guide on how to check every source AI gives you, and the same rules apply here. A cited report feels trustworthy, which is exactly why you should verify it.
From my own experience running websites and digital projects, the biggest win is the time shift. A competitor or topic overview that used to cost me an evening of open tabs now costs a coffee break plus twenty minutes of checking the sources. The checking part stays. Only the collecting part got fast.
A Simple Workflow for Students and Researchers
A workflow that works well in practice: start with one deep research run to map your topic. Then pull the real papers it points to and read them properly, using the approach from my guide on doing a literature review with AI. Finally, load your verified PDFs into a grounded tool like NotebookLM, which only answers from the documents you give it. My NotebookLM guide walks through that step.
Important tip: write your deep research prompt like a brief, not a question. Say what you need, for what purpose, in what format, and what to exclude. One detailed paragraph in produces a far better report than one short sentence.
Common Questions
Is AI deep research free?
Partly. Perplexity includes a limited number of Deep Research runs per day on its free plan. ChatGPT and Gemini include deep research with their paid plans, with smaller allowances on free tiers that change over time, so check the current limits on the official pages linked above.
Can I cite a deep research report in my thesis?
No. Treat it like a knowledgeable friend’s summary. Find the original sources it cites, read them, verify them, and cite those instead.
Which tool should I start with?
Perplexity, simply because you can try it today for free. If you already pay for ChatGPT or Gemini, use the one you have. For summarizing papers you have already collected, see my guide on summarizing research papers with AI.
Final Takeaway
AI deep research turns hours of collecting sources into minutes, and that changes how study and research feel day to day. But it moves the work, it does not remove it. Let the AI gather, then do the human part: read, question, and verify. Used that way, it is one of the most practical AI features you can add to your routine this year.
by admin | Jul 16, 2026 | Research & Productivity
Picture this: you ask an AI tool to help with your literature review, and it hands you a perfectly formatted reference, correct author names, a real-sounding journal, a plausible year. You paste it straight into your bibliography. There’s just one problem. That paper doesn’t exist.
This isn’t a rare glitch anymore. It’s become common enough that major journals, publishers, and research-integrity teams are now treating it as one of the biggest quiet risks in academic writing today. If you use AI for research, essays, or a thesis, this is worth five minutes of your time.
What is a hallucinated citation?
A hallucinated citation is a reference that an AI tool generates that looks completely real but doesn’t actually exist, or that misattributes real findings to the wrong paper. Researchers studying this problem call the worst examples “Frankenstein” citations, because they stitch together fragments of genuine papers (a real author, a real-sounding title, a real journal name) into something that was never actually published.
The dangerous part is that these fake references rarely look fake. They’re usually formatted correctly, attributed to real researchers, and dated plausibly. Unless you actually go and check, there’s often no obvious red flag.
How big is this problem, really?
Bigger than most people realize, and it’s growing fast. A Nature news feature published in April 2026 reported that tens of thousands of papers published in 2025 may contain invalid, AI-generated references.
A separate analysis is even more specific about the scale. Researchers led by Maxim Topaz at Columbia University audited nearly 2.5 million PubMed-indexed papers and published their findings as a letter to The Lancet in May 2026, reported in detail by Retraction Watch. They found that about 1 in every 277 papers published in the first seven weeks of 2026 referenced a paper that doesn’t exist. That’s a sharp jump from 1 in 458 in 2025, and 1 in 2,828 back in 2023, a roughly 12-fold increase in fabricated citations in just two years. The researchers traced the sharpest rise to mid-2024, right around when AI writing tools became widely used.
One more detail worth knowing if you’re writing any kind of review paper: the study found review articles had a fabrication rate 57% higher than other paper types, likely because they cite so many sources at once.
Why do AI tools make up references in the first place?
General-purpose AI chatbots like ChatGPT, Gemini, or Claude are built to predict the next most plausible piece of text, not to look things up in a verified database by default. When you ask one to “give me three sources on X,” it generates something that fits the pattern of a real citation, without necessarily checking whether that exact paper exists. It’s the same underlying issue behind AI giving wrong factual answers generally, which we cover in more depth in our guide to why AI sometimes gives wrong answers.
Researcher Maxim Topaz, who led the Lancet analysis, made an important point in his interview with Retraction Watch: most of the cases his team found weren’t researchers deliberately faking sources. 91% of the flagged papers had only one or two fabricated references, which he said are “likely honest mistakes by authors who used AI tools without verifying the output.” In other words, this usually isn’t dishonesty. It’s trust placed in a tool that was never designed to guarantee factual citations.
How to check every AI-generated citation
The good news is that verifying a citation only takes a minute or two once it’s a habit. Here’s a simple process:
- Search the exact paper title in quotation marks on Google Scholar or PubMed. If nothing comes up, that’s your first warning sign.
- Check for a DOI, and paste it into Crossref’s search tool to confirm it resolves to a real, matching paper.
- Open the actual source. Don’t just trust that the AI’s summary of a paper matches what the paper really says, skim the abstract yourself.
- Be extra careful with review articles and papers that cite many sources at once, since that’s exactly where this analysis found the highest fabrication rate.
Quick tip: if an AI tool gives you a citation you can’t verify within two minutes of searching, treat it as fake until proven otherwise, not the other way around.
From my own experience working on websites and digital tools, this is really the same instinct as checking a suspicious link before you click it. You don’t assume something is safe by default, you look for confirmation first. Citations deserve the same habit.
Tools that reduce this risk
Not all AI research tools carry the same risk. Some are built specifically to ground their answers in real, searchable sources rather than generating text freely. If you’re doing a literature review, our step-by-step guide to literature reviews with AI covers tools like Elicit and Semantic Scholar, which pull directly from real paper databases and show you the actual source, rather than describing one from memory. Similarly, our guide to AI tools for thesis writing and our walkthrough of summarizing research papers with AI both lean on tools that link back to the original document, so you can check the source yourself in one click.
Free citation managers like Zotero also help here, not because they use AI themselves, but because they store the actual paper alongside the reference, making it easy to double-check what you’re citing before you submit anything.
What this means if you’re writing a thesis, paper, or report
If you’re a student or researcher using AI to speed up your work, this isn’t a reason to stop. AI is genuinely useful for finding starting points, summarizing dense papers, and organizing your reading list, our beginner’s guide to AI covers the basics if you’re still getting comfortable with these tools. The real takeaway is simpler: treat every AI-generated citation as a draft that needs verifying, not a finished fact. That one habit is the difference between using AI well and ending up in a retraction story.
Common Questions
Can AI research tools like NotebookLM or Elicit still invent citations?
They’re much less likely to, because they’re designed to ground answers in the specific documents or database you give them rather than generating references from general knowledge. But no tool is risk-free, so it’s still worth spot-checking anything that goes into a formal paper.
Is using a fake AI-generated citation considered academic misconduct?
Opinions among researchers and publishers differ, and it depends on intent and how central the citation is to your argument. Most experts agree it’s treated far more seriously if you didn’t bother to check the source at all, so verifying every reference protects you either way.
How can I quickly tell if a citation is fake?
Search the exact title in quotation marks on Google Scholar or PubMed, and check the DOI on Crossref. If the paper doesn’t turn up, or the DOI doesn’t resolve to a matching title, treat it as unverified until you find it yourself.
Final takeaway
AI can genuinely speed up research, but it can also hand you a citation that looks completely real and isn’t. The fix isn’t complicated: search the title, check the DOI, and open the actual source before it goes anywhere near your bibliography. That one habit keeps AI a useful research assistant instead of a liability.
by admin | Jul 9, 2026 | Research & Productivity
Ask anyone who has written a thesis which part they underestimated, and you will usually get the same answer: the literature review. You start with one search, and two weeks later you have sixty open tabs, a folder full of unread PDFs, and no clear picture of the field.
AI tools can remove a lot of that pain if you point them at the right jobs. This guide walks through a five step literature review with AI, using free tools, and it stays honest about the parts you still need to do yourself.
Can you really do a literature review with AI?
Partly, yes. AI is genuinely good at three jobs here: finding papers that match your question even when you do not know the perfect keywords, summarizing individual papers quickly, and showing how papers connect to each other.
What it cannot do is judge research quality the way you can, decide why a gap in the field matters, or build your argument. AI models also make mistakes with total confidence. They sometimes invent references or misread a paper\u2019s findings, a problem we explained in AI hallucinations explained. So the workflow below uses AI for speed and keeps the judgement with you.
Someone close to me spends her days in PhD research on machine learning and medical imaging, so I have watched how fast a reading pile can grow. The researchers who cope are not the ones reading faster. They are the ones with a better system.
Step 1: Turn your topic into a real question
\u201cAI in healthcare\u201d is a topic. \u201cHow accurate are deep learning models at detecting brain tumours from MRI scans?\u201d is a question. Every step that follows works better when you start from a question, because modern research tools use semantic search. They match meaning, not just keywords.
Write your question down before you open any tool. If you cannot phrase it yet, that is useful information too. Spend an hour with a general overview or a textbook chapter first, then come back.
Step 2: Find papers with AI search tools
Three tools cover most of the discovery work:
- Elicit searches more than 138 million papers. You type your question and it returns a table of relevant papers with short summaries. The Basic plan is free, and it can import your library from Zotero.
- Semantic Scholar is a free academic search engine from the non-profit Allen Institute for AI. It indexes over 200 million papers and adds short AI generated summaries, called TLDRs, so you can screen results quickly.
- Research Rabbit maps papers visually. You start with one paper you already trust, and it shows similar, earlier, and later works, so you follow the citation trail instead of searching blind.
University libraries have started recommending these tools too. The University of Michigan Library keeps a practical guide on AI in literature reviews if you want a librarian\u2019s take on the same tools.
Tip: run the same question through two different tools. Each one searches differently, and the papers that appear in both lists are usually the ones to read first.
Step 3: Screen and organize what you find
You will collect far more papers than you need, so do not try to read them all. Screen each one by its abstract or TLDR and sort it into three piles: keep, maybe, and drop. Be ruthless with the drop pile.
For the keepers, use Zotero, a free and open source reference manager. It stores your citations, formats them in thousands of styles, and connects with Elicit and Research Rabbit. We covered where it fits in our guide to AI tools for thesis writing.
Step 4: Summarize and compare the papers
For every paper you kept, you want four things: the question it asked, the method it used, what it found, and its limitations. AI can speed this up a lot. Our guide on how to summarize research papers with AI shows practical prompts, and our NotebookLM and Elicit walkthrough covers tools that answer questions only from the sources you upload.
One warning from experience: AI extraction makes mistakes. It can misread a sample size or blur two findings together. Check every number and claim against the actual paper before it goes anywhere near your draft.
Step 5: Write the review yourself
Here is the part no tool can do. A literature review is not a list of summaries. It is an argument about the state of a field: what researchers agree on, where they clash, and which gap your work will fill. That structure has to come from your reading, so group your papers by theme or debate, not by author.
Two rules protect you here. First, verify that every reference exists and says what you claim, because AI generated citations are sometimes fake. Second, check your university\u2019s AI policy and disclose what you used. Most universities now allow AI for searching and summarizing but treat AI written text as misconduct.
From my own work with websites and online tools, the pattern is always the same. Tools that remove boring steps earn their place. Tools that promise to think for you cause trouble later.
Common Questions
Can AI write my literature review for me?
It can produce text that looks like one, but that is the trap. The references may not exist, the synthesis is shallow, and most universities treat submitting it as academic misconduct. Use AI to find, organize, and summarize. Write the argument yourself.
Are these AI research tools free?
Yes, for everything in this workflow. Semantic Scholar is completely free, Elicit has a free Basic plan, Research Rabbit lets you sign up free, and Zotero is free and open source.
How many papers should a literature review include?
It depends on your field and level. A bachelor\u2019s thesis might cover 20 to 40 papers, while a PhD literature review can pass 150. Your supervisor\u2019s guidance beats any general number, so ask early.
Final Takeaway
A literature review with AI is not about outsourcing the reading. It is about shrinking the boring parts: hunting for papers, formatting citations, and writing first pass summaries. Pick one question, run it through Elicit or Semantic Scholar this week, and save what you find into Zotero. The pile gets smaller, the map gets clearer, and the thinking stays yours.