The original post: /r/television by /u/RoryTate on 2024-07-12 13:12:47.
A lot of competing claims of “review bombing” and “review boosting” have appeared recently around the audience and critic scores on Rotten Tomatoes. This analysis attempts to prove or disprove these claims using 20+ years of data from the site. A public google spreadsheet is provided for full transparency into how these results were calculated. A more detailed breakdown of the data sourcing, methodology, etc, is included in a comment attached to this post.
TLDR; Analysis of the data shows a rampant and growing bias towards positive (“fresh”) scores being given by critics, resulting in significant review “boosting” of Tomatometer ratings over the last decade or more. A recent increase in audience scores since 2020 also suggests a problem with review “boosting” of audience ratings. At no time was any problem with review “bombing” seen in any of the scoring data. The primary explanation for the issue of positivity bias remains unidentified, but further detailed analysis of the data reasonably rules out several of the more common reasons for the issue.
Data
All audience and critic scores were first split into percentile groupings, with totals for each percentile calculated. A Bell Curve Distribution model was used for comparison. Normalized rating systems data predicts that the largest totals should occur closest to the average score, with significant drop offs on either side as the score percentiles move away from that midpoint. This “bell”-type distribution of scores can be seen reasonably well in the audience score percentile charts:
Note: the total number of critics reviewing a movie or show within a period of time was used in these charts to limit the entries to only the more well-known shows, by establishing a minimum of twenty critic reviews required within each time period to be included in the totals.
The audience scores appear to follow a nice, robust curve for the first twenty years. However, an unexpected increase is seen in the average score in the 2020-2023 data, alongside an unhealthy rise of scores in the 80 percentile range, which makes these specific ratings questionable. At no time does a widespread issue with “review bombing” appear anywhere in the observed data.
Note: To view any of the static charts in more detail, see the publicly available google spreadsheet.
Here are the Tomatometer (critic) score percentile charts for the same time periods, using identical criteria (requires 20+ reviews):
These Tomatometer percentile curves start off much more flat than the audience scores. In the first ten years the general “bell” shape is still somewhat visible, though it is biased toward higher scores. However, after 2010 the prevalence of scores in the 80-90 percentile range becomes highly distorted, until in 2020-2023 these two highest percentiles comprise the clear majority of scores for movies and shows. This indicates a serious and systemic problem with “review boosting” across all critic scores given on Rotten Tomatoes in the last decade.
In order to further explore this large trend towards positive (fresh) reviews, breakdowns of critic voting profiles and critic prolificity and impact are created in sheets 5 and 7. These analyses show remarkable changes in reviewing habits across all critics in the last ten or more years, leading to the observed “review boosting” behaviour. Also, deeper analysis of the data shows that the impact of reviews is not uniform across many variables, with strange variations being observed in critic activity across different time periods. The predicted behaviour of older critics creating more reviews – which should give them a greater impact on the site’s overall ratings – only holds true for critics who joined the site at its inception back in 2000-2004. Remarkably, critics who joined the site in 2015-2019 have the next highest level of impact by a significant margin, which cannot be explained with just the greater increase in the number of new critics added during that period of time. Critics who joined the site in the ten years from 2005 to 2014 are much less active than during any other period, which is an unexpected finding that deserves deeper investigation. Also, analysis of the data shows that Tomatometer scores are mostly formed by reviews from only the 350 most active critics. This small group is responsible for more than 50% of the total number of review entries on the site. This is a surprisingly small number of people to have control over this prominent and widely recognized scoring metric.
Lastly, five primary explanations exist for the observed positivity bias among Rotten Tomatoes-approved critics, which can be summed up as follows: money/access, social-proofing, aggregation methodology issues, activism/agenda, and new reviewer bias. The possibility of money/access boosting the scores is beyond the scope of this analysis. The possible impact of “social-proofing” (i.e. critics being influenced by the public nature of their “rotten”/“fresh” scores into potentially forgiving the poor quality of many movies and shows) is also beyond the scope of this analysis. However, the consequences of the aggregation methodology – because it only allows “fresh” or “rotten” choices when forming the Tomatometer scores – cannot be the cause of this problem due to the statistical impossibility of such. Meanwhile, the remaining two possible causes are explored in some detail in the “Critic Profile/Detailed Analysis” sections. To summarize the data findings for the two remaining explanations (see Sheet 7 for more details), both are found to likely have a measurable – but relatively small – impact on boosting the scores, and as such they are reasonably ruled out as the primary reason for the more than decade-long positivity bias in Tomatometer ratings.
Some quick statistical notes for those who are interested (everyone else can feel free to skip ahead to the “Conclusions” section). While the data in Table 4-8 may seem tautological to some (it can be summed up as “worse shows get fewer reviews”), this analysis actually suggests a different interpretation of the real problem therein: the “poor” quality shows and movies represented in these charts are not getting an adequate number of reviews to determine their actual score. Also, for the “Critic Profile Ratio” data in sheet 5, the purpose was to isolate reasonably active critics (>=20 reviews) voting within a set period of time, and to profile each individual critic’s behaviour when selecting “fresh” or “rotten” for that period. So in these charts the critics are compared to each other without any weight being given as to how many reviews they have done (beyond the base minimum of 20). However, in the “Ratios for 20xx Critics” data in sheet 7, the “time” axis was instead the release year of the movie/show, which abnegates the importance of looking at the data based on each individual critic’s pattern/profile. As a result all “fresh” and “rotten” votes are summed together under each “20xx Critic” banner, which is likely the more expected behaviour for comparing these critic review totals. Readers should be careful to view these two analyses differently, despite their apparent similarities. The “Critic Profile” charts in sheet 5 attempt to normalize the data so that more prolific critics do not affect the ratios, while the “Avg Critic Ratios” charts in sheet 7 tally the review totals without grouping by individual critic first.
Conclusions
Based on this data, it is reasonable to conclude that the audience score is a generally reliable standard when it comes to judging the quality of movies/shows/etc. However, the critic score is not a reliable standard of measure, and it appears that the problem with review boosting has become so widespread that Tomatometer scores are now practically worthless.
The full reason for this review boosting issue remains unidentified, but based on further analysis of the available data there is no evidence that it is due to aggregation methodology issues, political and social activism/agenda bias among critics, or the addition of large blocks of newly approved critics in recent years on Rotten Tomatoes.
Less frequently reviewed shows do not follow the same general trend as popular shows when it comes to the ratio of “fresh” to “rotten” ratings, and they are measurably more likely to receive lower scores, which are decided by a much smaller – and arguably inadequate – number of critics. This difference in critic behaviour could explain how the “review boosting” bias observed in this analysis has remained somewhat hidden from the general user base of Rotten Tomatoes over the years.
The majority (>50%) of reviews on Rotten Tomatoes can be attributed to a small group of only 350 of the most active critics on the site. Also, a larger majority (approx 75%) of the scores are attributable to critics who joined the site within two distinct time periods: 2000-2004 and 2015-2019.