All Devlog Posts
First Playtest Results cover

In September 2025 we ran our first structured playtest with people who were not the three of us. Twelve players, recruited through a roguelite Discord community, each playing for two to three hours. We watched six of the sessions over screen share; the other six submitted notes through a structured feedback form. This post is about what we heard, what surprised us, and what we changed.

We are writing this not because all the feedback was positive (it was not) but because we think documenting what early playtesting actually surfaces is more useful to other indie developers than another "our playtesters loved it" post.

How We Structured the Sessions

We gave players a brief written prompt before they started: play as you normally would, and describe out loud or in notes what you are thinking when you make decisions. We asked specifically about moments of confusion, moments where something felt unfair, and moments where the game surprised them in a way they found interesting.

We did not give tutorials or explain systems in advance. We have a tutorial sequence in the game but we wanted to know whether it was sufficient, and the only way to test that is to let players encounter it without pre-briefing. Three of the twelve players had significant confusion during the tutorial that we had to note.

We set aside all performance data from these sessions: we were not tracking run times or completion rates. This was a qualitative read, not a quantitative benchmark. We wanted to understand how players were thinking about and reacting to the game, not average run length.

What the Feedback Showed

The feedback organized itself naturally into three clusters.

Navigation clarity: seven of twelve players mentioned difficulty understanding where they were in the dungeon relative to where they needed to go. Players described getting lost between side rooms and the critical path, losing track of which exits they had already explored, and occasionally discovering that they had backtracked significantly without intending to. This was the most consistent piece of feedback across all twelve sessions.

Our reaction was divided. Omar and I had known that the navigation system was minimal: we display a minimap showing explored rooms but not the current objective path. We had justified this as a design choice. After the playtest, we revisited that justification and concluded we had been rationalizing something we had not fully designed. The navigation feedback was telling us that players were spending cognitive resources on orientation rather than on combat and decision-making. That is a resource allocation problem, not a difficulty calibration preference.

Combat readability: five players described moments where they took damage without understanding what had hit them. The cases were consistent enough to identify specific scenarios: enemies with ranged attacks placed in poorly lit rooms, hazard tiles that were visually similar to regular floor tiles, and fast-moving harassers whose attack animation was ambiguous. These are not systems problems; they are presentation problems. The underlying mechanics are not unfair but they are not communicating clearly.

Run variety perception: four players mentioned, in different ways, that after their third or fourth run the dungeon was starting to feel familiar. Two of them used the word "same," which was concerning because our layout uniqueness metrics at that point showed high structural variety. The gap between our measured variety and the player's perceived variety is a design problem we did not anticipate as clearly as we should have.

What the Feedback Did Not Surface

Some things we expected to be problems were not mentioned. The enemy behavior system, which we had worried might produce confusing or apparently random enemy actions, generated no consistent negative feedback. Players either did not notice the trait variation in enemy behavior or did not find it problematic. One player specifically mentioned that enemies felt "unpredictable in a good way," which is approximately what we were aiming for.

Combat difficulty balance was mentioned by only two players, and in opposite directions: one found early rooms too easy, one found them too hard. Given that we had not calibrated difficulty for these specific player skill levels, that spread felt like a roughly correct starting point rather than a calibration failure.

What We Changed

In response to the navigation feedback, we added a directional indicator system that marks the critical path exit in each explored room with a subtle visual indicator. Players can now always tell which direction progress lies. We intentionally kept the indicator understated: a small floor marking rather than a compass overlay or waypoint arrow. We do not want to remove the navigational texture of the dungeon, just ensure players are not allocating attention to recovering from genuine disorientation.

For combat readability, we adjusted the visual contrast for hazard tiles across all dungeon depth tiers and added a brief flash effect to ranged attack projectiles in low-light rooms. We also extended the attack wind-up animation for harassment-type enemies by about 150 milliseconds. These are small changes but the playtest feedback was specific enough that we could address the root cases directly.

The run variety perception problem is harder to fix with a patch. We have two hypotheses about its cause. First, we may be generating layouts that are structurally varied but spatially similar: different room graphs with similar visual density and room size distributions. Second, our early dungeon tiers may not be varied enough in their enemy ecology, making the first few runs feel more consistent than later runs do. We are investigating both and have not yet settled on the intervention. This will take several more iteration cycles.

What We Would Do Differently

We should have run this playtest earlier. Not because we were not ready; we were ready enough. But we delayed because we wanted to feel more confident in the systems before exposing them to external eyes. That confidence-seeking delayed feedback we could have used three months earlier.

The navigation problem in particular is one that was visible in our own internal playtesting but that we had rationalized away because we knew the dungeon layout and were not affected by orientation confusion. An external player brought an honest encounter with the system that we could not simulate for ourselves. We knew we needed that; we should have gotten it sooner.

We are scheduling the next structured playtest for eight to ten weeks out, after integrating the current changes and building out the dungeon depth beyond what we tested this round. More on that when we have results worth reporting.