How Do University Students Use AI Chatbots? Reviewing Anthropic’s Education Report
Anthropic’s analysis of 540,000 Claude conversations offers a rare view of how university users delegate academic work to a chatbot—and raises difficult questions about higher-order thinking, assessment, and cognitive offloading.
A rare look at actual chatbot use
While preparing a perspective paper on how we should teach students in the age of AI chatbots, I came across Anthropic’s education report, published on 9 April 2025, and read it immediately.
The report analyzes 540,000 Claude conversations associated with university use. Anthropic used CLIO, an AI analysis system designed to protect privacy, so that people did not directly read the original chat transcripts.
This matters because there are many survey studies asking students, “How do you use chatbots?” Large-scale analyses of the chats themselves are much harder to conduct. It would be interesting to see OpenAI conduct a similar study, although whether it will is another question.
What students use Claude for
The largest category, accounting for 39.3%, involved generating questions from educational materials, revising essays, or summarizing content. A further 33.5% involved explanations or solutions for assignments. The remaining uses were more specialized.
- Educational-material-based question generation, essay revision, and summarization: 39.3%
- Assignment explanations or solutions: 33.5%
- Data analysis and visualization: 11.0%
- Research design and tool development: 6.5%
- Diagram creation: 3.2%
- Translation or review: 2.4%
Which fields use it most
Business, health care, and the humanities used Claude less than might be expected from their student populations. Science and mathematics, and especially computer science, used it more heavily. Computer science students represented 5.4% of the population but accounted for 38.6% of chats.
Social sciences and history, engineering, and education showed usage roughly proportional to their population shares.
The relatively low share for health care was unsurprising to me. In conversations with medical students, I have found them fairly conservative about these tools, and some medical schools and hospitals block access to chatbots on campus, partly because of patient-information protection.
Four ways of using the chatbot
The report classifies chatbot use along two axes: the desired output, either problem-solving or output generation; and the interaction style, either direct, task-commanding use or collaborative, conversational use.
The four resulting categories were used at broadly similar rates, ranging from 23% to 29%. This seems like a useful 2×2 framework.
Examples of more positive uses included explaining philosophical concepts and theories, generating learning materials for chemistry study, and explaining anatomy and physiology concepts relevant to solving an assignment.
Examples of more problematic uses included asking for answers to machine-learning multiple-choice questions, asking for answers to an English test, and rewriting marketing or management-related text so that it would not be detected by plagiarism tools.
Patterns by discipline
Science and mathematics students often asked for step-by-step problem-solving processes. Collaborative interaction was more prevalent in computer science, engineering, and science/mathematics.
In the humanities, business, and health care, direct and collaborative use were roughly evenly split. Education had the highest share of output-generation use, at 74.4%. However, the report notes that educational-material generation for other disciplines may have been misclassified into education.
Bloom’s taxonomy and the concern about higher-order work
In the analysis based on Bloom’s taxonomy, Creating and Analyzing were overwhelmingly common. Output generation is often classified as Creating, while problem-solving is often classified as Analyzing.
This does not mean that, when a chatbot performs these tasks, students are not doing them themselves. Still, the inverted-pyramid pattern is concerning because it suggests that students may be relying on AI for higher-order thinking.
Bloom’s taxonomy can make an analysis feel richer with just a spoonful added. I may have to use it myself in the future. But there is also an important limitation: the taxonomy was designed around students’ cognitive activity, so applying it to AI is not straightforward, particularly for categories such as Remembering.
Limitations of the report
The report has substantial limitations. Early adopters are likely overrepresented, especially because people who use Claude rather than ChatGPT are likely to be a very small group. Students also use tools other than chatbots for academic work.
CLIO’s classifications may be wrong. Although the chats were classified as student conversations, they could in fact have been from instructors or related staff. Because of the privacy policy, the analysis covered only 18 days of usage history.
The study analyzes what people delegate to a chatbot, but cannot directly answer how they study. Interdisciplinary content may have different usage patterns, yet the research uses a single classification system.
Questions for teaching and assessment
As students increasingly delegate higher-order cognitive work to AI, how can they develop critical thinking and metacognition? How should assignments and assessment change?
Anthropic is working with some universities to provide a Learning Mode chatbot service. My own view is that students should be taught ways of using these tools through a framework, alongside evidence-based information about the harms of cognitive offloading. The message should not simply be “Do not use it,” but rather, “Research suggests it can be harmful in these particular ways.”
Assignments should also be designed on the assumption that AI will be used. They could combine elements such as providing prompts and critically evaluating AI outputs. Oral assessment may need to carry more weight as well.