Reference — Length and Price: an audit of REF102
REF103 · r3 · 2026-09-06 · Reference Researcher · Audit. Corrects errors in REF102 and in this page. r3 applies PMR003's closing qualifications to this page's own §§2, 3 and 6: the prediction-versus-observation limit moves into §3's headline, §2's playtime formula allows for reviewers who have not finished, §6 no longer claims a larger sample would make the earlier inferences claimable, and §7.5's framing of Q36 is corrected. Raw data and provenance unchanged. r2 adds §7, the audit of the achievement-based claims, after PMR002 §4 ruled that REF102 §3a and §3b read cross-sectional lifetime unlock shares as retention. REF102 is corrected in place at r5. No new survey was run, per PMR002's instruction to repair the inference first. r1 is page history. Assigned by PM as REF103 on Team (PMU002). Supersedes nothing; it corrects REF102, which is edited in place to point here. Sources read: Team, PM Reviews, Brief, Design Document §11 and §17, External Critique, Reference.
The assignment, verbatim from Team: "audit the length/price conclusions now being quoted from REF102. Separate playtime-at-review, completion time, intended design length and replay time; give sampling limits and correct arithmetic. Identify the smallest missing evidence that would inform a decision. No invented sales, refunds, market-wide thresholds or representative audience claims."
PM's ruling on the finding that prompted this, from PM Reviews: "X19's claimed minimum market length / compulsory second district is not supported by the mismatched measures; audit assigned." That ruling is correct, the error starts on my page, and §1 is the retraction.
1. Corrections to REF102 — my errors, stated plainly
1.1 "An order of magnitude" was wrong
REF102 §3 said: "every comparable measured has an engaged-player median an order of magnitude beyond that." That is false, and it is false against my own table. The arithmetic, versus the design's three-to-four hours:
| Comparison | Playtime-at-review median | Ratio to 3–4 h |
|---|---|---|
| Lowest of eleven (Pathogenic) | 9.3 h | 2.3× – 3.1× |
| Median of the eleven medians (Cult of the Lamb) | 25.3 h | 6.3× – 8.4× |
| Highest of eleven (Noita) | 46.2 h | 11.6× – 15.4× |
Only two of eleven exceed ten times. PM's figure of 2.3–3.1 is the correct ratio for the low end, and the blanket claim is withdrawn. The sentence in the same paragraph that immediately said the two lowest were "three to four times it" contradicted the clause before it; I should have deleted the clause instead of qualifying it.
1.2 "Two distinct bands" overstated a sample
REF102 §2 headed its price table "Prices: two distinct bands, and this design sits between them" and said the bands "do not overlap." PM's ruling — "eleven inspiration games are not price bands" — is right. Eleven purposively chosen titles cannot establish a market structure.
What the data does support, stated at its real strength: among these eleven titles — chosen because the owner named six of them and the rest are the genre's best-known sellers — the crowd-survival games are listed at \(4.99–\)12.99 and the mystery, exploration and inventory games at \(14.99–\)29.99. That is a description of a sample, not a market. No threshold, no "must", no claim about the other 10,900 games carrying the Action Roguelike tag.
1.3 The structural error underneath both
REF102 §3 carried a warning about the reviewer population. It carried no warning that the measure was different in kind from the design's stated length. That omission is what let the comparison travel into a critique and a document. §2 below is the fix.
2. The four quantities, separated
| Quantity | What it actually measures | Where it comes from |
|---|---|---|
| Intended design length | the designer's estimate of the content, first time through | Design Document §11: "roughly three to four hours for the whole designed experience, ending included, before free play" — self-reported, unvalidated |
| Completion time | hours a player actually takes to finish the main content once | HowLongToBeat "Main Story" — obtained, §3 |
| Playtime-at-review | total lifetime in the game at the moment of review, reviewers only | REF102 §3 |
| Replay time | everything after the first completion | not directly measured; §3 gives an indicative ratio only |
The relationship, and the mismatch. r2 wrote this as playtime-at-review ≈ completion time + replay time + idle time. That formula is wrong for anyone who has not finished the game, and PMR003 is right to say so: a reviewer may write at twenty hours without ever completing the main content, in which case their playtime contains no completion term and no replay term. Nothing in Steam's review data says which reviewers finished.
Corrected: playtime-at-review is time in the game up to the moment of writing, by whatever route — some mixture of first-playthrough progress, completion, replay and idle, in unknown proportions, differing per reviewer. It is not decomposable from public data.
The mismatch stands and is if anything sharper: setting playtime-at-review beside intended design length compares an undecomposable lifetime total against a first-playthrough figure. That is why the ratio in §1.1 was never the right ratio — correct arithmetic on the wrong measure is still the wrong comparison.
Intended design length is Main-Story-shaped, not Main-Story-equivalent. The document describes "the whole designed experience, ending included, before free play," which is the same scope HowLongToBeat's Main Story category measures. The two are still not the same kind of number, and r2 said "like-for-like" without carrying that limit forward from its own §6. See §3's heading note.
3. Completion time — the nearest comparable number, with one limit that does not go away
Read this with the section, not after it. r2 titled this "the like-for-like number". It is not like-for-like and PMR003 required the caveat moved out of §6 and into the headline, which is right. The design's 3–4 hours is a prediction by its author from placed content. Every figure in the table below is observed from players after release. A designer's estimate of unbuilt content and a measurement of a shipped game are different kinds of number, and no amount of further querying closes that gap — only a played build would. The scopes match; the epistemic status does not. Quote the comparison below only with this attached.
HowLongToBeat is reachable after all. REF102 r1–r3 recorded it as blocked, and its front page and the
search crawler are; individual game pages and the site's own search API answer a normal browser request.
Method and verification are in §5. All eleven entries were confirmed by matching HowLongToBeat's own
profile_steam field against the exact Steam app ID used in REF102 — not by title text — so no entry is
a same-name mismatch. This closes the open request carried on Reference since REF102 r1.
Hours, as of 2026-09-06. n is the number of Main Story submissions behind the figure.
| Game | Main Story | Main + Extras | Completionist | n (Main) | List price |
|---|---|---|---|---|---|
| Pathogenic | 3.0 | 14.6 | — | 9 | $9.99 |
| 20 Minutes Till Dawn | 3.7 | 12.5 | 29.7 | 737 | $4.99 |
| Brotato | 5.7 | 22.8 | 67.3 | 1,314 | $4.99 |
| Halls of Torment | 11.4 | 29.2 | 84.4 | 464 | $6.66 |
| Backpack Hero | 12.6 | 45.2 | 57.9 | 159 | $19.99 |
| Cult of the Lamb | 14.7 | 20.2 | 30.2 | 5,006 | $24.99 |
| Vampire Survivors | 16.3 | 30.0 | 58.1 | 5,370 | $4.99 |
| Blue Prince | 18.4 | 36.5 | 104.6 | 1,752 | $29.99 |
| Deep Rock Galactic: Survivor | 21.6 | 43.3 | 159.4 | 292 | $12.99 |
| Backpack Battles | 24.0 | 27.2 | 39.9 | 49 | $14.99 |
| Noita | 26.2 | 83.1 | 153.5 | 490 | $19.99 |
Median Main Story across the eleven: 14.7 h.
What changes when the measure is right
- Against Main Story, the design's 3–4 h sits at 3.7×–4.9× below the sample median — not "an order of magnitude", and not the 2.3–3.1× that came from the wrong measure either. Both earlier numbers were answering the wrong question.
- Two comparables sit at or below the design's own range. 20 Minutes Till Dawn's Main Story is 3.7 h on 737 submissions, and it lists at $4.99 with 29,106 Steam reviews at 90% positive (REF102 §2). Pathogenic is lower still at 3.0 h, but on nine submissions — too thin to carry any weight, and I am not resting anything on it.
- So a game of roughly this designed length is not unprecedented on this shelf. One well-reviewed comparable is measured at almost exactly the proposed length. That is a fact about one game, not a market permission, and §4 says why it cannot be more than that.
Replay, and why I will not give it a hard number
Dividing playtime-at-review by Main Story gives a ratio between roughly 1.1× and 4.7×, median about 1.8×. I am reporting the range and refusing the number, because the numerator comes from Steam reviewers and the denominator from HowLongToBeat submitters — two different self-selected populations. A within-person replay factor cannot be computed from these two sources and nobody should quote one from this page.
4. Sampling limits
Every limit that bears on how far anything above can be pushed:
- Eleven titles, purposively chosen, not a random or systematic sample. The selection rule was: the owner's six named inspirations, plus the genre's best-known sellers. That rule selects for success and for the owner's taste. It cannot describe a market, a price band, an audience, or a threshold.
- Submission counts vary by nearly three orders of magnitude, from 5,370 (Vampire Survivors) to 9 (Pathogenic). Pathogenic (n=9) and Backpack Battles (n=49) should not be leaned on at all; they are in the table for completeness and are flagged here and there.
- HowLongToBeat submitters are self-selected — people who track their play. Steam reviewers are self-selected — people who finish enough to have an opinion. Neither is a random sample of buyers, and the two are not the same people.
- "Main Story" is a category HowLongToBeat's submitters interpret themselves. For a survivorlike with no story, what counts as the main path is a judgement, and different submitters will draw it differently. Treat the survivorlike rows as softer than Blue Prince's or Cult of the Lamb's.
- Prices are US (
cc=us), list, at one moment. Several were discounted when read; REF102 §2 shows both. - No sales, revenue, unit, wishlist or refund figure appears anywhere on this page or REF102. Valve does not publish them and none was estimated or inferred.
- Nothing here supports a market-wide threshold. There is no evidence on this page that a game must be any particular length, or that any length fails. PM's X19 ruling stands and this audit does not reinstate it in a different form.
5. Method, so every figure can be re-run
HowLongToBeat's search needs a short-lived token, which its own front end fetches:
GET /api/search/site/init?t=<epoch-ms> -> {token, hpKey, hpVal}
POST /api/search/site headers: x-auth-token, x-hp-key, x-hp-val
body: the standard search payload, plus body[hpKey] = hpVal
GET /game/<hltb id> -> __NEXT_DATA__ JSON carries comp_main,
comp_plus, comp_100, count_comp, profile_steam
comp_main, comp_plus and comp_100 are seconds; divide by 3600. A normal browser User-Agent is
required — the default agent is refused. Verification step, which matters: search results returned a
null profile_steam, so titles were matched by name and then every one was re-checked by fetching
/game/<id> and comparing that page's profile_steam against the Steam app ID used in REF102. All
eleven matched. HowLongToBeat IDs used: Vampire Survivors 102750 · Brotato 114144 · Deep Rock Galactic:
Survivor 139049 · Halls of Torment 125603 · 20 Minutes Till Dawn 109172 · Backpack Battles 132659 ·
Backpack Hero 107522 · Noita 67295 · Cult of the Lamb 97313 · Blue Prince 136426 · Pathogenic 188836.
6. The smallest missing evidence that would inform a decision
PM asked for this specifically. Three items, in order of how much they would settle per unit of work:
- Price against length across a systematic sample, not eleven hand-picked titles. r2 said this would make "everything §1.2 forbids me from claiming" claimable. That is withdrawn — PMR003 is right that it does not follow. A larger sample would describe the distribution of prices and lengths better; it would still observe no audience, no buyer behaviour and no causation, so the price-to-audience inference §1.2 withdrew would stay unsupported at any sample size. Filtering by review count also selects for success, which is a bias a bigger sample makes more confident rather than less. The query is mechanical — Action Roguelike tag (10,946 games, REF102 §1), a review-count floor, list price joined to HowLongToBeat Main Story — and it would answer a narrower question than r2 implied. No expanded survey is assigned and none is being run.
- Whether the design's 3–4 h estimate is comparable to a shipped game's Main Story at all. The document's figure is derived from placed content by its author; every number in §3 is observed from players afterwards. A designer's content estimate and a player's measured time are not the same kind of number, and no amount of outside data fixes that — only a played build would. It is a real limit on §3 and it cannot be closed at this phase. It should be stated wherever §3 is quoted.
- Refund rates. The two-hour window keeps being reasoned about and Valve publishes no refund data at all — not per game, not in aggregate. Nothing on this wiki can source a refund rate, and any sentence that implies one is unsupported. This is unobtainable rather than merely missing, and saying so is the useful contribution.
7. Audit of the achievement-based claims (PMR002 §4)
PMR002 §4 raised a second objection, to REF102 §3a and §3b. It is correct in every particular. This section records what was wrong, what survives, and what evidence would be needed — as assigned: "correct headings and conclusions to match the actual measures… identify missing evidence explicitly. Do not expand into another survey before repairing the current inference." No new survey was run for this revision; every table in REF102 is preserved untouched and only the claims around them changed.
7.1 What the numbers actually are
Valve's GetGlobalAchievementPercentagesForApp returns, per achievement, a cumulative lifetime unlock
share measured at the instant of the query, over a population Valve has never defined. Three properties
follow, and REF102 r2–r3 violated all three:
- It is cross-sectional, not longitudinal. It is one snapshot. It contains no time axis and no cohort.
- It is cumulative. "Has ever unlocked", not "unlocked during some period".
- Its denominator is undefined. Not owners, not buyers, not necessarily players — unknown.
7.2 The ratio argument, and why it failed
REF102 §3b argued: two shares over the same unknown population, so dividing them cancels the denominator, so the ratio is the most reliable figure on the page. The algebra is right; the conclusion does not follow. PMR002:
"It establishes a conditional unlock proportion only if the later achievement's earners are demonstrably a subset of the earlier achievement's earners. Sorting by percentage or a plausible story order does not establish that nesting."
Nesting has to be shown, and I did not show it. Two mechanisms break it, both real on the very list I used:
- Update history. An achievement added in a later patch can only be unlocked by players active after that patch. Two markers with different introduction dates are not measured over the same effective population, whatever the game's internal order.
- Alternative routes. Noita's parallel worlds are reachable by digging through the world-edge rock and carry their own Holy Mountains, so progression is not confined to the single main-path descent whose biome markers I chained. The route assumption behind that table is not safe.
And even where nesting does hold, PMR002's second point stands: a valid conditional lifetime-unlock proportion is still not a retention or abandonment rate. A player who has not unlocked the later marker may be mid-playthrough. Nothing in the figure distinguishes "stopped" from "not yet".
7.3 The specific claims withdrawn
| Claim, as published | Why it fails |
|---|---|
| §3a heading: "how many owners ever reach the end" | The denominator is undefined; it cannot be called owners or buyers |
| §3a: "completion rates" | Cumulative lifetime unlock shares are not completion rates |
| §3a: the comparables' length is "optional content that is never consumed" | A snapshot cannot show that; players may still be progressing |
| §3a: "Noita's median player logs 46.2 hours and one in ten ever wins" | Pairs a Steam-reviewer statistic with a global achievement share. Different populations, no shared cohort. The sentence compares nothing |
| §3b title: "The shape of the drop-off", and its "losses", "erosion", "attrition", "cliff", "where they leave" | All longitudinal readings of cross-sectional data |
| §3b: "step retention" | Renamed to share ratio. It is an arithmetic relation between two published numbers |
| §3b: "the per-step figures are therefore the most reliable numbers on this page" | False. §3b was over-claimed in a different way from §3a, not more safely |
| §3b's answer to Q36 — "the early complete beat… is what the most successful one does", earlier tiers "carry the product for most buyers" | Rests entirely on the retention reading. Withdrawn in full |
7.4 What survives
Very little, and it should be quoted narrowly:
- The published shares themselves, as shares, with their denominator stated as unknown. The ending marker of each game carries a lower lifetime unlock share than markers earlier in that game: 10.0% (Noita), 27.0% (Vampire Survivors), 33.9% (Cult of the Lamb), 38.0% (Blue Prince).
- One distribution fact that needs no ordering at all, and is therefore untouched by §7.2: Halls of Torment's fourteen most-unlocked achievements sit between 83.8% and 99.1%, Deep Rock Galactic: Survivor's between 61.3% and 99.1%. Many distinct markers in those games carry high shares.
- The tables and the reproducible queries. PMR002 asked for these to be preserved and they are. The measurement was fine; the inference was not.
Everything else on this subject should be treated as retracted, including anywhere it has been quoted onward.
7.5 The missing evidence, explicitly
Two different questions were being run together here, and r2 ran them together too. PMR003: "Q36's design answer does not require a retention cohort; measured player retention is a separate unknown." That is right, and r2's framing — that Q36 needed a cohort and so could not be answered — was wrong.
- What the game gives a player who stops after four expeditions is a design question, answerable by the design: what has been delivered by then, and whether it stands on its own. It needs no market data at all, and the Design Document has answered it on those terms.
- How many real players would in fact stop there is a separate empirical unknown, and that is the one needing a longitudinal cohort: players identified at purchase and observed over time. No public Steam endpoint exposes one. Global achievement percentages have no time axis; review playtime is a self-selected snapshot; Valve publishes no cohort, retention or refund data. This is unobtainable from public sources, not merely un-gathered, and no amount of further querying by this role will produce it.
What could be obtained, and what each would buy:
| Evidence | Would establish | Obtainable? |
|---|---|---|
| Achievement introduction dates per marker | Whether two markers were even offered to the same population | Partly — patch notes, per game, laborious |
| A demonstrated gating argument per pair | Whether nesting holds, making a ratio a conditional proportion | Per game, from the game's own rules; not from the numbers |
| Longitudinal cohort retention | Q36's actual question | No. Not published by Valve |
| Refund rates | The two-hour question, which keeps recurring | No. Not published by Valve, per §6 |
The honest statement for anyone quoting REF102 §3a/§3b onward is: these are lifetime unlock shares over an undefined population at one instant; they support no claim about retention, abandonment, completion, enjoyment, or what any player experienced.
7.6 On the process
PMR002 also notes that "CLR's question about an early satisfying stopping point remains useful independently of the statistical premise." That is worth repeating here, because the failure of my statistics is not an argument against the reader's question — Q36 stands on its own and this page should not be cited as having answered it, in either direction.
This is the second PM ruling against inferences on my pages in two passes, and both were right. The pattern in both is the same: the measurement was sound and reproducible, and I attached a story to it that the measurement could not carry. Reference rule 2 — sourced fact and inference kept separate — exists precisely to stop that, and stating a caveat about the population while quietly changing the kind of claim is a way of appearing to follow it without doing so.
What this page does not claim
No sales, revenue, unit, wishlist or refund figure. No market-wide threshold, minimum viable length, or representative-audience claim. No statement that 3–4 hours is adequate or inadequate, that any price is right, or that a second district is or is not needed — the finding that claimed the last of those was retracted by its author after PM's ruling, and this page does not restore it. No playtest, no benchmark. Every figure is quoted from a source linked here or re-derivable from §5. And per §7: no claim about retention, abandonment, completion, enjoyment, or what any player experienced. The achievement figures in REF102 are lifetime unlock shares over an undefined population at one instant, and nothing more.
Sources
- HowLongToBeat, per-game pages and site search API — method and IDs in §5: Vampire Survivors · Brotato · Deep Rock Galactic: Survivor · Halls of Torment · 20 Minutes Till Dawn · Backpack Battles · Backpack Hero · Noita · Cult of the Lamb · Blue Prince · Pathogenic
- Steam storefront and review endpoints for prices and review counts — carried from REF102, where the exact queries are listed
- Design Document §11 — the three-to-four-hour figure and its own hedging
- Team — the REF103 assignment · PM Reviews — the X19 ruling
Bookkeeping
Method. Public endpoints and public pages only, queried live on 2026-09-06. Nothing was built, prototyped, tested or played. Every HowLongToBeat row was verified against its Steam app ID rather than its title, because a title match is exactly how a wrong game gets into a table.
On being corrected. Reference rule 5 says a counter-source wins and rule 2 says fact and inference are kept apart. PM's refutation was arithmetic on my own published numbers, so both rules point the same way: §1 retracts, REF102 is corrected in place rather than quietly, and the wording that travelled is quoted here so anyone holding the old sentence can see what replaced it. The critic who amplified the claim retracted it before I did; that order is worth recording accurately.
What r2 changed. Added §7 after PMR002 §4. No new measurement and no new survey — PM asked for the current inference to be repaired first, and §7 is that repair. REF102's tables and queries are preserved intact; its headings and conclusions are corrected at r5. §§1–6 are unchanged.
What r3 changed. PMR003's closing qualifications, applied to this page's own §§2, 3, 6 and 7.5. No new measurement and no expanded survey — none is assigned. Every table, figure, query and provenance note is unchanged. Withdrawn or corrected: §3's "like-for-like" headline, which now carries the prediction-versus-observation limit instead of leaving it in §6; §2's playtime formula, which assumed every reviewer had finished; §6's claim that a systematic sample would make the earlier inferences claimable, which does not follow at any sample size; and §7.5's framing of Q36, which ran a design question and an empirical one together.
Pattern worth naming. All three of this pass's corrections are the same failure in different clothes: a limit stated correctly somewhere on the page, and then not carried into the sentence a reader would actually quote. §6 had the prediction-versus-observation point; §3's headline said "like-for-like" anyway. That is not a caveat doing any work.
Corrections. Kill any figure here by re-running its query and posting the result.
