Our first structured playtest in September 2025 surfaced a finding that troubled us: four out of twelve players described the game as feeling repetitive after their third or fourth run, even though our generation metrics at the time showed high structural variety across runs. Players were reporting sameness that our instruments were not detecting.
That gap between measured variety and perceived variety sent us back to the question of what we are actually measuring and whether our metrics correspond to the player experience we care about. This post describes what we found and what we changed.
The Easy Metrics Are Not Enough
The most natural metric for procedural run variety is layout uniqueness: compare the room graph of run A to run B and measure structural overlap. We were running this check and it showed high uniqueness scores. Room graphs across a sample of one hundred consecutive runs had low pairwise similarity by graph isomorphism comparison. The dungeon layouts were genuinely structurally different.
The problem is that structural uniqueness is necessary but not sufficient for perceived variety. Two dungeons can have different room graphs but feel similar because they share the same visual density, the same average room size, the same distribution of hazard types, and the same patterns of enemy encounter. The player is not experiencing the room graph; they are experiencing the traversal of those rooms, the fights inside them, and the choices available at each point.
We were measuring the right thing to know that our grammar was not degenerating into repetitive structure. We were not measuring the things that determine whether a player feels like they are exploring something new.
The Five Metrics We Now Track
Room role sequence distinctiveness. We hash the sequence of room roles along the critical path (Entry, Combat, Transition, Resource, etc.) and measure how many distinct sequences appear in a run sample. This captures whether the narrative rhythm of the dungeon varies across runs, not just whether the room graph varies. A dungeon where the critical path is always Combat, Combat, Resource, Boss Approach, Boss produces a consistent rhythm regardless of how spatially different the room layouts are.
Ecological encounter overlap. We track which ecological compositions (predator count, harasser count, role ratios) appear in the combat rooms of each run and measure overlap between runs at the same depth tier. This tells us whether players are seeing the same types of fights at each depth level across multiple runs. Two runs can have different room graphs but if combat rooms at depth two always have a predator-anchor composition, the encounters feel the same even if the rooms look different.
Spatial density distribution variance. We measure the distribution of walkable tile density across rooms within a run and compare the distribution shape across runs. This captures whether dungeons vary in their spatial character, not just their topology. A run that consistently generates medium-density rooms across all depth tiers will feel spatially monotonous even if the graph structure changes.
Behavior trait set overlap. We measure which behavior traits appear in enemy instances across runs and calculate pairwise overlap in the trait vocabulary a player encounters. If the same subset of traits appears in nearly every run, the enemies will behave similarly regardless of how their stats are scaled or which room they appear in.
Decision density per run segment. We count the number of meaningful player choice points per room (exit choices, resource collection decisions, encounter routing options) and measure the variance of this count across runs and across depth tiers within runs. This is the hardest metric to define cleanly because "meaningful choice" requires a definition. We currently use a proxy: any moment where the player pauses for more than three seconds in a non-combat context. This is imperfect but it captures the surface of player deliberation reasonably well.
What We Found When We Applied These Metrics
The ecological encounter overlap metric identified the problem our playtesters were describing. At depth one and depth two, our encounter compositions had low variance: the ecology system was consistently drawing from a small subset of compatible unit pools for those depth tiers. Technically this was correct behavior, because we had marked fewer units as depth-one-compatible during the early design period. But the practical effect was that the first two thirds of most runs had a similar encounter vocabulary, creating the "sameness" feeling even in players who went deeper and found more variety later.
The room role sequence metric also showed lower distinctiveness than we expected. Not because the grammar was failing to produce varied sequences, but because some sequences are more likely than others under our current constraint weights. We had a run of about three months where the probability distribution over critical path sequences was not well calibrated: certain sequences had high probability and others were theoretically reachable but practically rare. Players were not noticing that the sequences were varied because the most common sequences dominated their experience.
What We Changed
We expanded the depth-one and depth-two ecology unit pools, adding more role-compatible units at low depth tiers to increase the encounter variety early in runs. The effect was measurable in the ecological overlap metric within two weeks of the change: pairwise composition overlap dropped by about a third at depth-one combat rooms across the test suite.
For the critical path sequence distribution, we adjusted the constraint weights to pull probability mass away from the most common sequences and distribute it more evenly across the range of valid sequences. This required careful testing because reducing the weight of common sequences can increase the frequency of edge cases in the grammar, some of which produce less-pleasant room arrangements. We found a calibration point that improved sequence variety without increasing edge case frequency significantly.
We have not solved the variety perception problem completely. The spatial density distribution metric still shows lower variance than we would like at mid-dungeon depths. Rooms in the middle of the dungeon (depth two and three) tend toward medium density because our grammar's room size constraints are tightest there. We have not yet found a grammar adjustment that increases spatial variety at those depths without creating navigation problems. It remains an open calibration question.
Why This Is Hard to Get Right
The fundamental difficulty in measuring run variety is that the metric has to correspond to player cognition, not to a formal property of the generated artifact. Players are not comparing room graphs; they are forming impressions about whether the game is surprising them. Those impressions depend on memory, on context, on expectation. A room that was novel in run one is familiar by run four not because the room changed but because the player has encoded it.
We cannot build a metric that directly captures player memory and expectation formation. We can build proxies that we believe correlate with those things, calibrate them against playtest observations, and track whether improving the metrics improves player reports of variety. That is the process we are running, and it requires ongoing playtest validation to know whether the correlation is holding as we make changes.
What we can say confidently is that structural layout metrics alone are not sufficient proxies for perceived run variety, and that measuring variety at the encounter and rhythm layers provides signal that structural metrics miss. The playtesters were right: the game was less varied than the layout metrics suggested. The newer metrics helped us understand why.