Addressing grade inflation
(because I mentioned it yesterday)
A fundamental problem with how we grade now is that we are trying to use one score, a letter grade, for two purposes---ranking students and describing how well they’ve mastered the course material. You can’t cram that much information into one score, and you shouldn’t try. I think that a lot of the confusion around grading stems from people’s failure to distinguish these two goals.
If we only report letter grades and believe that “A” means “top 10%” and that it also means “mastered the course material”, then we apparently believe that only 10% of the students can master the course material. This might be true in some courses, but it’s a silly restriction.
I suggest we use letter grades to indicate level of mastery of the material, and numerical scores to convey rankings.
Universities should have some courses in which any student who works hard enough will be able to master the material. For such a course, we might assign many “A” grades. Many introductory courses could be like that1. We can also have courses that require unusual effort and talent just to get by. In these, a student might be ranked in the top 10% and still not get an “A”.
I prefer grading schemes that allow us to convey both ranking and mastery of material. My favorite approach is to have instructors assign only letter grades, but have transcripts indicate the rankings of the students who received that grade. For example, it could say something like “B+, ranks 10-15 out of 50”. This is the solution recommended by Yale’s Committee on Trust in Higher Education2. It is easy to implement: instructors assign grades as usual and the registrar computes the ranks.
Another reasonable variation is to ask the instructor to report both a letter grade indicating level of mastery of material and a ranking in the class. Both would appear on the transcript. My complaint with this approach is that a total ranking can magnify small differences that break ties. It might be that there’s no real difference between students who are consecutive in ranking. I prefer to allow ties by reporting rankings in ranges. One could achieve that by allowing an instructor to report both letter grades and ranges of rank. So, some students could get “A, 1-2 out of 50” while others get “A, 3-5 out of 50”. This has the advantage of allowing an instructor to make finer distinctions between students while also allowing them to group students they think are equivalent. But, this is probably too much work for too little advantage over just assigning grades.
This does leave the question of what should become of grade point averages. To start, we could report two: one based on letter grades and one based on rankings in classes. To average ranges of rankings, we could average the middles of the ranges. If half the students in a class get an “A”, then an “A” will correspond to a score of 75%. In a class where only 10% of the students get an “A”, it would receive a score of 95%. This gives students who want high averages an incentive to take courses in which high grades are rare, and it gives professors incentives to use the whole grading scale when it is appropriate. When computing “Latin honors”, like Summa Cum Laude, we should use averages of course rankings rather than numerical averages of letter grades.
Some additional considerations
Grade inflation isn’t just a problem in education. Ratings throughout digital commerce have been compressed at the high end. If you give an Uber driver any rating less than 5 stars, Uber wants to know what they did wrong. They don’t treat 3 stars out of 5 as average. In fact, they drop drivers with average ratings below 4.5.3
If different schools start grading very differently, it will be harder to compare students across schools. Those looking at transcripts will have to weight grades by school and year. This will be more difficult for smaller institutions that see fewer transcripts.
Many years ago, Brendan Hassett pointed out to me that one university had a lower grading scale for very introductory math courses than for the more advanced courses available to students who had taken advanced math in high school. This discrepancy imposed a grade penalty for students from poorly resourced high schools. Providing rankings by class would even this out.
We sometimes teach courses that I describe as X for future presidents, where X is something like Math, Data Science, or AI. These courses teach people who aren’t going to go into a field enough that they can ask good questions and figure out how to make good decisions. Such courses to not have to be very difficult to be valuable.
