Methodology
How the VGC Score works
A review score should describe the game you can buy today, not the one that launched months ago. VideoGamesCritic treats a game's rating as a living estimate that updates as the game does. No editorial judgment, no frozen verdicts: just a published, reproducible formula over public data.
1. Every Steam review is trust-weighted
Not all reviews carry the same information. Using fields Steam publishes with every review, each one gets a weight between 0.04 and 1:
- Playtime at review. Reviews written with under 2 hours played (inside Steam's refund window) are heavily downweighted (10–25%). This neutralizes "thumbs down, then refund" reviews and drive-by ratings. Weight ramps linearly to 100% at 10 hours.
- How the game was obtained. Non-Steam key activations count 60%; copies received for free count 40% of that. Review copies and bulk-key giveaways shouldn't drown out paying players.
- Review bombs. Steam's own off-topic review-bomb detection stays enabled, so flagged periods are excluded at the source.
2. The score re-anchors at significant patches
Patch notes from the developer's Steam announcements are classified by significance: Steam's own patchnotes tag, title keywords (major update vs. hotfix), body length, and substance keywords (balance, rework, performance, new content). A crash-fix hotfix does not reset anything; a balance overhaul does.
Reviews written after the latest significant patch describe the current game, so they form the score's primary evidence. If no significant patch is known, the window falls back to the last 30 days.
3. Steam approval is calibrated to a press-comparable scale
A Steam positive-review rate is an approval measure, not a quality score: the median Steam game sits around 80–85% positive, so raw percentages read absurdly high next to press scores. The VGC Score maps approval onto the familiar 0–100 quality scale before blending: 95% positive ≈ 89.5, 80% positive ≈ 74, 50% positive ≈ 52. A game at 96% on Steam and 85 on Metacritic isn't a contradiction; they're different scales, and the VGC Score reconciles them.
4. Everything blends in a Bayesian model
The VGC Score is a weighted blend built from pseudo-review counts:
- Players hold the majority, by rule. This is the core difference from Metacritic and OpenCritic, which are press-first aggregates. Here, one professional review carries the evidence of ~2 player reviews (capped at 100 press reviews), and press evidence as a whole is hard-capped at half the player evidence, so players always control at least two-thirds of the verdict. Press informs the score; it cannot outvote the people who bought the game. Before players have shown up, press gets only a small fixed voice, so brand-new games start near the prior and get player-verified within days. When OpenCritic hasn't covered a game, Metacritic's critic score stands in.
- Tracked creators count like press. Reviews from the YouTube creators we track carry the same per-review weight as press and share the same cap against player evidence. Creators who score games contribute their scores; verdict reviewers like Skill Up contribute their recommend/skip calls, pooled into an approval rate and calibrated on the same curve as Steam approval. A wave of creator pans registers even when the outlet never printed a number.
- Launch-era evidence decays with age. Press reviews and Metacritic user aggregates describe the launch build. After a 1-year grace period their weight halves every 3 years, while current player evidence keeps full strength. This is the "living" in living score: No Man's Sky launched to a 49 here and sits in the 80s today, because a decade of updates outvotes a 2016 review, something a frozen aggregate can never show. Every game also stores its launch quality (press + first 12 weeks of Steam sentiment), so the "+34 since launch" story is explicit.
- All-time Steam sentiment is context, not verdict: capped at 2,000 reviews and halved. That's enough voice that a stable classic's lifetime record isn't outvoted by one quiet month, but the post-patch window still dominates when a game genuinely changes. And the window only carries full weight when a fresh significant patch anchors it, meaning a new build is being judged. With no recent patch, the last 30 days are just chatter about an unchanged game and count at a fraction, so an Ori-tier classic isn't dragged under its lifetime verdict by one quiet month.
- Metacritic user ratings are the counterweight to Steam's self-selection: Steam reviewers bought the game, while Metacritic collects console players and the broader audience. Ratings are unverified, so each counts half a review (capped at 1,500, ignored below 30 ratings), but a thousand-strong verdict genuinely moves the score. This is how a game at 88% on Steam and 5.9 user score lands in the mid-70s instead of the 80s.
- Post-patch, trust-weighted reviews are the evidence, scaled by the real review volume in the window (capped at 3,000).
- A "typical game" prior (quality 72, weight 50 pseudo-reviews) keeps a tiny game with forty glowing reviews from outranking established greats. Evidence washes it out; hype alone doesn't.
- Elite scores must be earned at scale. A 99% positive rate from 5,000 reviews and an 89% from 180,000 are not the same kind of evidence: small games are reviewed almost exclusively by the self-selected fanbase most likely to love them, while a game bought by hundreds of thousands faces the whole spectrum of players. So the score range above 88 unlocks progressively with total player volume (Steam lifetime reviews plus Metacritic user ratings): no headroom at 1,000 reviews, the full scale at 100,000. A small gem can absolutely reach 88 ("excellent"), but calling something an all-time great is a claim about breadth of acclaim, and the ceiling rises as the audience grows. Two refinements keep this honest: console-only games get their off-Steam ratings counted at a Steam-equivalent multiplier that ramps in with volume (so 27,000 Metacritic ratings for a Nintendo exclusive count like the millions of players they represent, while a title with 1,700 ratings cannot borrow that multiplier), and games no press outlet ever reviewed only unlock half the elite range, because universal acclaim means critics and players agree. On the other side of the same coin, truly massive consensus earns extra weight: past 100,000 lifetime reviews, the evidence cap on all-time sentiment grows with volume instead of treating 855,000 reviews like 35,000.
The practical effect: a handful of post-patch reviews barely moves the score, but a large wave of them dominates it. A game that launched rough and got genuinely fixed climbs; a game that broke after an update falls. Each score also carries a confidence level based on total evidence, and a trend badge comparing the last four weeks of reviews against the twelve before them (volume-gated so quiet weeks don't read as signal).
Early Access games carry no VGC Score. A work in progress isn't the shipped game, so scoring it would grade a draft. Instead their pages show the signals that matter for a work in progress: whether players recommend the current build, how sentiment is trending, how long it's been in Early Access, and whether the developer is still shipping updates. The moment 1.0 ships, the game is scored completely fresh: the scoring window, launch reception, and lifetime sentiment all restart at the full release, so years of Early Access reviews (glowing or rough) never bias the verdict on the game that actually shipped.
5. We rate the reviewers, too
Every press outlet and creator we track gets a trust score: a measure of how well their review scores align with what players think of the same games. Press records come from OpenCritic and outlets' own sites; creator records also come straight from their YouTube channels, where each review's stated verdict is read from the title or description. Only explicit calls count: a creator who never states a buy/skip verdict is not graded, because interpreting their opinion for them would stop being arithmetic. It is not an opinion about anyone's writing or motives; it is arithmetic over their published track record.
- Fair ground truth. An outlet's verdict is compared against player sentiment contemporaneous with the review: the 12 weeks of Steam reviews starting when it was published, calibrated to the same 0–100 scale. A launch review is judged against launch sentiment, a re-review after the big patch against post-patch sentiment. A reviewer who praised a great launch isn't punished because the publisher ran the game into the ground years later (Overwatch), and one who panned a broken launch isn't punished by the fix. Reviews published before a game existed on Steam have no observable player signal for that build; those are excluded entirely rather than compared against a different era's players.
- Scales are decoded, not punished. Reviewers have never used the full 0–100 range: a "7/10" is press-speak for mediocre, and even weak games get a 6. So before measuring anything, each outlet's scores are normalized within its own scoring history and mapped onto the player-quality distribution. What's judged is whether knowing the outlet's verdict tells you what players will think, not whether it uses the same numbers players do. The raw habit is still published as its own stat ("hypes +9"), it just isn't double-punished.
- Alignment (40%). Rank correlation: does the outlet sort games the way players do, its best and worst calls matching players' best and worst.
- Decoded accuracy (30%). After scale decoding, the average distance from player quality. Low residual error means the outlet is genuinely informative once you know how to read it. This is the headline stat on every outlet card: "typically within 6 pts of players".
- Player-side (30%). The controversy test. On games where press consensus and players sharply disagreed (15+ points, the Veilguard cases), whose side was this reviewer on? Following the press herd against players drags trust down; calling it like players did raises it. Every outlet page spells this record out: on the games where press and players sharply split, how many times this reviewer scored it the players' way.
- Verdict-only reviewers count too. Many creators (Skill Up and others) publish recommendations rather than numbers. They're rated primarily on call accuracy (50%): when they say buy, do players end up satisfied; when they say skip, do players agree? That's blended with alignment (25%) and the player-side controversy test (25%), and labeled "verdict-based". Tiered verdicts that aggregators convert into numbers (a "Buy" stored as 100) are detected and rated as verdicts too, because a two-value scale is a recommendation, not a score, and treating it as a score would manufacture fake "scored 100, players felt 60" disagreements.
- Recency. Review weight halves every two years. Trust measures who you can rely on now: an outlet that drifted from its audience is judged mostly on its recent record, and one that course-corrected gets credit for it.
- Shrinkage. The blend is pulled toward a neutral prior until an outlet accumulates real history, and outlets under 20 usable reviews are listed as provisional. Six lucky reviews can't top the leaderboard.
- Uncertainty is published, not hidden. Every trust score ships with a bootstrap 95% confidence interval (the ± next to the number). While an outlet's interval overlaps its neighbors', its rank is marked approximate (≈); we will not sell noise as a ranking. As the review corpus grows the intervals tighten and ranks firm up on their own.
Some games have no gradeable player verdict, and we refuse to invent one. When a massive audience is split near the middle (PUBG at 1.0: half a million owners, half of them thumbs-down, all of them playing every night), neither a 95 nor a 40 is provably wrong, so the game grades nobody. When player sources contradict each other about the verdict itself (Metacritic users bombed Genshin Impact while 400k PlayStation raters call it excellent), the ground truth is contested and the game is skipped. And a few thousand self-selected ratings never speak for a mass-market audience critics broadly endorsed. Cultural phenomena get the same humility at the other end of the scale: when a game's real audience is in the hundreds of millions (Minecraft), the couple million people who bothered to leave a store rating are a contrarian-skewed sliver, so anything short of clear acclaim or a clear pan from that pool grades nobody. Reviewers are only graded on games where players actually reached a verdict.
Only reviews of games with at least 500 Steam reviews count. Each outlet page shows the scatter of its scores against players, its controversy record, and its biggest disagreements: the receipts behind every number. The same math powers the press-vs-creators comparison on the trust leaderboard: two camps, one standard. Creators are YouTube or Twitch channels plus a named allowlist (Angry Joe, Jimquisition, Skill Up, and others). A one-person site stays press until it is named; we do not guess from bylines.
The test this algorithm has to pass: an outlet being rated should have nothing to argue with. It is judged only on games it chose to review, against what players thought at the time of the review (contemporaneous sentiment: a fair review of a broken launch is never punished by the later fix, and a fair review of a great launch is never punished by later mismanagement), with its scoring dialect decoded rather than penalized, with old reviews fading, with small samples pulled to neutral instead of amplified, and with the remaining uncertainty printed next to the score. There is no editorial input anywhere in the pipeline; the number is a pure, reproducible function of the outlet's published reviews and public player sentiment. Disagree with your rating? The way up is written into the math: call games the way players end up seeing them.
6. Full transparency, always
Every game page shows the exact components (positive rate and pseudo-review weight of each source, the scoring window, the sample sizes) so any number can be checked. An aggregate you can't audit is just another opinion.
The "what players say" panel on each game page is built the same way: aspect mentions (combat, performance, grind, monetization, and so on) are counted across a sample of recent and most-helpful Steam reviews, split by review verdict, and the sentiment chart annotates dips and jumps with the patches and complaint themes from that exact window. It is plain keyword counting over published reviews. No text on this site is machine-generated.
Data sources: Steam reviews, review histograms, and developer announcements; OpenCritic press scores and per-outlet reviews; Metacritic critic and user scores, including per-platform critic scores; PlayStation and Xbox store ratings; IGDB metadata.
7. Ports are not the Steam build
Steam reviews describe the PC game. A Switch or last-gen port can be a different, worse product, and the VGC Score on that platform must not inherit the Steam number as if every conversion were perfect.
- Canonical score stays Steam-first. The homepage, unfiltered browse, and the hero number on a game page answer "is this a good game?" from the living PC (or console-exclusive) evidence.
- Platform browse uses that platform's evidence. Metacritic critic scores for Switch, PlayStation, or Xbox, plus that platform's store ratings when they exist. Steam never counts as Switch evidence. Press-only is enough: awful ports usually show up in platform critic scores first.
- Thin evidence stays honest. If a port has fewer than four platform critic reviews and no platform player ratings, we keep the canonical score and mark it PC ratings rather than inventing a penalty or pretending the Steam number is a Switch verdict.
8. How the platform holders are doing
Switch, PlayStation, and Xbox browse each include two holder scores: how that publisher did on the previous box (Switch, PS4, Xbox One) versus the current one (Switch 2, PS5, Xbox Series). A game is filed by its original release date, so God of War (2018) stays PS4-era even if Steam dates the PC port 2022, and remasters (Spider-Man Remastered, Days Gone Remastered, The Last of Us Part I) stay with the original box rather than padding PS5. True cross-gen launches like Ragnarok and Miles Morales still count as this generation. It covers games that holder published (Xbox Game Studios, PlayStation Studios, Nintendo, Bethesda, and their labels) on that hardware. A studio they later bought does not pull in old work-for-hire: Stick of Truth stays Ubisoft, not Microsoft, even though Obsidian made it. Sony-published third parties (Stellar Blade, a Shift Up game) drop off once they ship on Xbox or Switch; Bloodborne still counts because it never did. First-party studios still count after ports, so Indiana Jones remains a Microsoft game after a PlayStation release. Last-gen (Xbox 360, PS3, Wii) does not. A Valve PC game that later appeared on Xbox does not.
The numbers are Bayesian blends of those games' canonical VGC Scores, notable releases only (press coverage or a full-priced game with real player volume), weighted by log audience size, with the same typical-game prior so a handful of 95s cannot mint a fake 94. One Steam capsule cycles beside the scores, mixing highs from both gens so Ori still appears on Xbox. Hover or tap a generation to see every game in that blend. Vintage remasters stay out of the pager. Bloodborne and Bayonetta 3 still count because Sony and Nintendo published them. There is no PC holder panel.