LLMs, Testing, Grade inflation, and Accommodations
Like every professor, I’ve been worrying about how AI systems will force us to change our teaching practices. We also worry about how we will handle grade inflation and what I’ll now call accommodation inflation. Unfortunately, how we deal with each of these is going to complicate how we deal with the others. LLMs are so good at quantitative homework that, instead of evaluating students through homework, we will have to rely much more on tests. Pressure to decrease grade inflation will greatly increase the stress those tests create. Students who have trouble with tests will need more accommodations.
Below, I’ll describe these problems in more depth and the logistical problems they will create. I’ll also mourn what we’re losing.
I hope to one day be able to write about the bigger question of how advances in AI should change what we teach.
Replacing Homework with Tests
LLMs have gotten very good at programming and doing math. They can probably solve all the homework problems we might want to assign in a class. This makes it unreasonable to use homework to evaluate and grade students in our classes. We still want them to do homework, because it’s the best way for them to learn some of the things we want to teach. But it is simply too easy now for students to ask LLMs for the answers. The natural solution is to tell students that homework is just for their training, to stop grading it, and to grade them through tests or oral presentations. My biggest problem with this is that the skills I usually want students to acquire are difficult to measure by testing. I don’t have as much experience with oral exams, but I do know that it’s logistically difficult to conduct them in a large class. Let’s think about what will happen if we only evaluate students through tests.
Our classes will have many more tests, and many classes will have tests that didn’t before. We will need to set aside more time for these tests. At Yale right now, the final exam period is already very crowded, and students who take many classes with final exams often encounter timing conflicts between them. Current policy at Yale recognizes that three exams in a row is too many. We are going to have to make the exam period longer. Some students who have disabilities are afforded twice as much time to take their exams and tests. Those students will face even more scheduling conflicts. Many professors who’ve thought this through are now giving short quizzes many times per semester. It is even harder to give extra time for these, because the rooms aren’t available before or after classes. We are going to need to build extra time into the schedule for tests.
Grade Inflation
At the same time, universities are starting to reduce grade inflation1 (Harvard) or considering it2 (Yale, 2). Grades have risen steadily over the decades since I’ve been a professor. I’ve contributed to this. Way back in 2013, a large Yale College Faculty meeting discussed the distribution of grades and suggested ways of changing our system. This meeting was the first time that I learned how high the grades were in typical Yale courses, and that the grades I assigned were much lower. Around 20% of the grades I assigned in my core Computer Science course were in the C/D/F range. I didn’t want students to suffer just because they had me as a professor or just because they were majoring in Computer Science. So, I raised my grades.
One advantage of grade inflation for professors, and the system in general, is that it makes small differences disappear. We might want to know about small differences in students. But we’d rather mask out the small differences in the tests they take. When students miss tests, either because of illness or conflicts with other tests, we are supposed to offer them a make-up test. This test can’t be the same as the original, because students might have heard about the questions. So we create fresh make-up tests and pretend that they have the same difficulty as the original. They don’t. They could be easier or more difficult, but they will almost never be the same. I try to adjust my grading to compensate for my estimate of each test’s difficulty, but that’s guesswork. If the exact scores on tests didn’t matter so much, I wouldn’t worry about it. But more detailed grading will create more pressure to be fair and make testing much more difficult.
Accommodation Inflation
The accommodations that irritate professors the most are exceptions that deans can grant to excuse missing tests or assignments. At Yale, these are now called “Dean’s Extensions”. They used to be “Dean’s Excuses.” In theory, these are a good idea. Sometimes a student has an emergency or serious illness that causes them to miss a test or prevents them from turning in a problem set on time. Rather than making the professors try to figure out which student’s claims are legitimate, a dean who lives in the student’s dorm and who knows their history makes the decision. The problem is that the number of such extensions that are granted has increased dramatically. Tenfold is a reasonable guess. This is a problem at many schools. It got much worse during COVID. And that’s OK with me. I’d rather not have a COVID-positive student coughing in my class, or even one with the flu. But it’s made the job of testing difficult. In a class of 20 students, at least one will need a make-up exam. In a class of 100, some will miss the make-up and need make-ups for the make-ups. As these drag on, it becomes difficult to figure out who will create these exams and who will grade them. Sometimes these make-ups happen after a professor has left. Maybe we ask the next person teaching a course to do it, if there is such a person the following semester or year. But the material won’t be identical. Creating a good test takes a lot of work, and it’s a big ask of someone who didn’t even teach the course.
The academic accommodations that have received the most attention are alternate testing arrangements for students with disabilities. Some students need technological support during tests, some need extra time, and some need isolation or an unusual location. Yale has an office and facilities to provide that support. As the number of tests increases, we are going to need a lot more of these facilities. Students who need extra time for tests are going to need a lot more time. We need to start thinking about where to find it.
Increased testing and more granular grading will increase scrutiny of the accommodations for students with disabilities. Fairness dictates that we ensure that accommodations are carefully tailored to the students and the tests. No matter what we do, there will be suspicion of the students whose disabilities are not obvious: pretty much everyone would like extra time to finish tests. We have to start dealing with this now before the issues become a political football. Recent press reports on the increasing fraction of students who receive accommodations because of disabilities suggest it may already be too late3 (3,4,5,6).
What we lose
Grades
I’m not a fan of grades as a motivation for students. I would rather that students learn because they are interested. The school I attended through 8th grade didn’t have grades, and the experience shaped me profoundly. In high school, the pressure to get good grades decreased my intellectual motivation. College was much better, because grade pressure was less intense than in my high school, and because I felt that the grades in the classes I was taking purely for intellectual curiosity didn’t really matter.
But sometimes we need grades to choose which student is right for an opportunity, like who to hire as a research assistant, and when we need to recommend a student try something ambitious, like pursuing a Ph.D. Our current grading system makes it difficult to distinguish between students, and it makes it difficult for them to figure out where they stand relative to their peers. Assigning grades by adding test scores seems like a solution. Test scores make it easy to assign as many grades of each letter as we want. But tests don’t always provide good measures of what we really want our students to learn.
Homework
Working through homework problems is the best way that I know to learn material. Struggling to figure out how to apply new techniques to solve difficult problems forces one to understand the limits and capabilities of those techniques. This struggle helps students grow.
I’ve put a lot of effort into dreaming up homework problems that will help my students learn. I like to assign problems that can only be solved after serious thought and a non-trivial insight. No one should finish my problem sets in an hour. I want students to practice thinking hard and understand the problems so well that they can think about them while walking to class, while taking a shower, while waiting in line, or any other time they might be bored. I want them to feel excited when they solve a problem.
The students who can routinely solve these problems are the ones I encourage to pursue research.
It’s difficult to find problems like this, and it’s especially difficult to find new ones each year. I used to look for them in every paper I read and every talk I attended. My notes are full of pointers to lemmas that might make for good homework problems. It wasn’t unusual for me to spend 10 hours designing a problem set. I will probably be relieved that I won’t have to do that anymore. Now I can just assign old problems as homework, and hope that students work through them. But we’ll be losing a lot.
I should explain that I don’t think all classes should assign problems like this. I think that students need many different types of educational experiences and many different types of classes. They should have classes where they collaborate and classes where they work on their own. They should have classes where they learn a little in depth, and classes where they cover a lot. They should have classes that they take just to learn something they’d never encounter otherwise. I like teaching the classes that prepare them for graduate school.
Tests
It’s perverse to respond to the strength of LLMs at doing homework by evaluating students through tests, because LLMs are even better at the questions we can pose on tests! In the age of LLMs that can program and do math, tests reveal even less of what we want from our students. I doubt that tests will be great predictors of whether students will be able to do great research. But we are heading towards a system in which we only grade students by their performance on tests, and in which small differences in performance on those tests could result in big differences in grades.
Tests can be an indicator of talent. There are brilliant students who do very well on tests. But there are many brilliant students who don’t score the best on tests, and I worry that it will be harder to discover them. If I’m not confident that I can identify the best students in my classes, I’ll be even less sure of my ranking of the rest.
My skills
As a student, I spent a lot of time watching my teachers and thinking about how they taught. When I became a professor, I tried to incorporate the best of what I’d seen, to the extent that it was coherent and fit with my personality. I’ve spent decades becoming the best teacher I can be. Advances in AI are going to force me to change and rethink how to teach. Technological advances often eliminate entire categories of work, so I should not be surprised. But I‘ll mourn the irrelevance of hard-earned skills, and I will feel bad for my students until I figure out how to teach well in this new reality.
Caveats
Technological advances have caused educators to panic before, and I haven’t been sympathetic. But LLMs feel different. When people said that typing is worse than writing by hand, that reading on paper is better than reading on a screen, or that taking notes by hand is better than taking notes on a computer, I’ve been skeptical. And when I’ve read the studies that supported these assertions, I’ve been underwhelmed. Some have the same intellectual validity as a hypothetical experiment in which we randomly assign students to take a test with their left or right hand, observe that the students in the right-hand group perform better, and then conclude that all students should write with their right hands.
I always tell my classes that different students learn differently, and that their job in college is to figure out how they learn best. Neither students nor educators should assume that what works best for most students will work best for all.
I don’t know how we will resolve these problems. I prefer to try to motivate students to learn and to optimize my course for the students who want to learn, as opposed to devoting all my efforts to grading and preventing cheating. Distrust is a poor basis for the relationship I want to establish. If it weren’t for the pressure to decrease grades, I’d probably respond to LLMs by simply not grading. Or, in an ideal world, I’d teach only small classes in which I got to know every student well. But we have too many students for that to be a viable solution.
Stories
My thinking is of course informed by my experiences. Here are a few that are relevant to this blog post, in no particular order.
There was a time when faculty were concerned about all the screens in their classrooms. Many students told us they were typing notes, and I know some were, but faculty worried about what they might be watching and whether they were distracting their fellow students. I won’t write now about the low quality of the studies I read on the topic. Instead, I’ll tell you what I did. I asked a teaching fellow to hang out in the back of the class and tell me what the students were looking at. They told me that most of the students were using their laptops to take notes, to look up things that I mentioned in class, or follow along with the lecture notes I’d provided. I adopted a classroom technology policy that amounted to “be courteous of those around you.”
One of my favorite professors when I was an undergraduate was Richard Beals. He gave very difficult exams, and explained that a score of 50% would earn a B. He thought there was no point in putting questions on exams that he expected us to answer, so every question would be difficult. He didn’t want any student to leave early, so some questions would be very difficult. But we could earn an A without being perfect. I found that very freeing. If perfection is the standard, then the task is too easy. When I started teaching, I gave exams like that. I learned that what I found freeing, other students found traumatizing. Too many students were crying during my exams. One of my best students told me that when they looked at the exam they weren’t sure if they would be able to solve any of the problems. I eventually started giving easier exams.
One year while grading my exams, I noticed two strange answers that were so similar that I was sure one student had copied from another. To check if my suspicion was credible, I looked at a picture I had taken of the exam room. Those students were sitting so far from each other that copying would have been impossible. It is more likely that they studied together, and developed the same strange way of thinking about the material.
Many years ago, the Dean of Yale College discovered a collection of class materials in a bathroom during exams. So, she suggested that we not let students leave class to use the bathroom during tests and exams, or that we have them work on the exams in sections, with bathroom breaks between. I wasn’t enthusiastic about this idea, in part because I posed difficult exam questions and wanted to give students time to think about them. In the first test I administered after the dean’s announcement, a student in class blew their nose, and developed an unfortunately severe nosebleed. I decided that student should be excused to use the bathroom.
Keeping the dean’s concerns in mind, I occasionally followed students who left exams to use the bathroom. I tried to do it covertly, leaving a few minutes after they did. I always felt like an ass. What I heard told me that many of those students really did need to use the bathroom, and that keeping them in the classroom could have been disastrous. I decided that my default should be to trust my students, and that I don’t want to try to enforce draconian testing rules just to prevent cheating. I try to establish a good relationship with my students that will help motivate them, and focusing my pedagogy on catching cheaters instead of teaching interested students would make it hard to establish the relationship I want.
https://www.nytimes.com/2026/03/02/us/colleges-students-disabilities-enrollment.html,
https://www.thetimes.com/us/news-today/article/40-percent-stanford-undergraduates-claim-disabled-sw99r3k8c,
https://www.theatlantic.com/magazine/2026/01/elite-university-student-accommodation/684946/,
https://stanforddaily.com/2026/04/09/the-real-reason-students-disabled/
