Yale’s Committee on Teaching, Learning, and Advising is working on a proposal to improve grading and the FAS/SEAS Faculty Senate just approved a resolution on the matter1. So, I’m writing more about my preferred solution, which is to use ridit scores2.
I’ll begin with the principles I’d like a solution to satisfy:
I want to provide incentives instead of imposing requirements. I want to avoid a solution in which instructors are required to give a certain number of grades of each type.
All instructors and courses should be equal.
Students should not be given reason to prefer courses with generous grading policies.
Instructors should not be under pressure to grade generously or harshly.
It should be easy to implement.
The solution I know that satisfies these requirements does not change the grades used, but rather how they are aggregated and reported. Each student will receive both a grade and a score. Instructors assign letter grades, with pluses and minuses as they see fit. These grades are then automatically converted into scores. Whenever students are compared, they will be compared by scores rather than by grades. Classical GPAs based on grades would be abandoned and would be replaced with the average of these scores, which we might call the “Yale-GPA”.
The scores will be numbers between 0 and 100, and the average score in every course will be 503. The amount of score that an instructor has to distribute among the students in their class will be fixed and equal in every class. So, no instructor will be viewed as an easy or tough grader, at least when we look at the scores.
The scores will be determined by how many students in the class received each grade. Instructors will not have to do any work to compute these scores. For each letter grade, we will compute the percentage range of students who received that grade. The score is the middle of that range. Here are a few examples of how this would be computed:
If 20% of the students in a class received a grade of B and 20% received higher grades, then the students who received a B are those between the 60th and 80th percentiles. All those students receive a score of (60+80)/2 = 70.
If the bottom 10% of the students in a class receive a grade of D, their score will be (0+10)/2 = 5.
If 50% of the students in a class receive a grade of A, they will all receive a score of (50+100)/2 = 75.
If all the students in a course receive an A, then they all get a score of 50. They get the same score if they all receive a grade of C.
This scheme guarantees that the average score for every class will be 50, up to the rounding errors that occur when we compute middles of ranges.
Reactions I expect
We need to think about how people might react to this. Here are some reactions I anticipate.
An instructor might decide to give every student in their class an A. That means that all those students will receive a score of 50. The top students in the class will be unhappy, because they would have received a higher score if more grades were used.
Some instructor might think that just one student in their class is amazing. They might decide to give that student the only A, so as to give that student a high score. If there were 10 students, that score would be 95.
The average Yale-GPA will be 50. We will now learn who the average students are and discover that average students are actually pretty smart.
A Yale-GPA of 75 will be very high.
There will be much more variance in scores. So, a low score in a few classes actually won’t hurt one’s GPA all that much.
There will be stress in small courses because the average score will be 50. If all of the students are good, I expect instructors to respond by assigning everyone the same grade. This won’t be horrible, because 50 will be the average grade and amazing students will have opportunities to get high grades in other courses. I believe that the benefits of being in a small class outweigh the risk in grades.
Variations
Some people attach meanings to numbers, like “60 and below means failing.” To make those people more comfortable, we could rescale the numbers from 0 to 100 to a range of 60 to 100, by the formula 60 + 0.4 * score.
We could get rid of the letters A, B, etc. We could replace them by terms like “excellent”, “good”, or “satisfactory”4. As long as the terms are ordered, we can convert them to a score. This would have the advantage that it would be harder to convert them into a GPA on the standard 0-4 scale. If we report both grades and scores on student transcripts, some will try to convert those grades to the 0-4 GPA scale. And, we’ll want to discourage that so that our students don’t suffer in comparison to those from other universities.
We would also have to decide what information to report on a transcript. If we are reporting grades for each course, we could report the percentage range of students who received that grade. That would be more informative than just reporting the middle of that range. We should also report the number of students in the class. And, wherever we might provide a GPA, we should provide the new average score.
Alternative formulation
Here’s another way to understand these scores. Imagine an instructor who has 10 students in their class, and that they have actually ranked all of them. We could imagine letting them report that ranking, or the induced scores which would be 5, 15, 25, ..., 95. If the instructor thinks that a group of students are actually equivalent, they could put those students in the same bucket and average their scores. For example, they might decide that the 2nd and 3rd students from the bottom are equivalent, and give them both a score of 20. Or if the top 3 students are all equivalent, they would all get a score of 85, which is the average of 75, 85 and 95. Grouping the students preserves the averages.
Open problems
If we implement this system, instructors are going to change how they grade. They will probably do this in reaction to the feedback they get every year. Can we predict where they will eventually wind up? That is, under a reasonable model of student and instructor utility functions, can we predict the eventual equilibrium? If so, would it help to provide it to everyone at the start?
How will this break? What have I missed?
Here are my previous posts on this topic:
Up to some small rounding errors.
This is like the “change the currency” proposal made to the Yale College Faculty in 2013
