All Devlog Posts
Three People, One Dungeon Engine cover

One of the first decisions you make as a small studio building a procedural game is where to draw the custom-versus-off-the-shelf line. Write your own engine or use an existing one? Write your own procedural framework or build on a library? Write your own behavior tree runtime or use a plugin?

There is no universally right answer, and our answer has evolved as we have understood the actual constraints better. This post describes where we landed and why, including some places where we got it wrong initially.

The Base Engine: Godot

We build Ruinveil in Godot 4. The choice was pragmatic: we were three people with different background engines (one of us had primarily Unity experience, one had Unreal experience, and I had worked mostly with custom engines in previous roles). Godot is MIT-licensed, has no royalty structure, and has a scripting layer that is accessible without being overly opinionated about how you structure game logic. The editor tooling is good enough for the tile-based dungeon visualization work we need without requiring custom editor extensions for the common cases.

The limitation we ran into with Godot for our use case is in runtime asset streaming. Our dungeons generate at run start and room assets need to load quickly without hitching the frame budget. Godot 4's ResourceLoader threading model required non-obvious workarounds to get reliable async loading behavior. We spent about three weeks on this during early production and it remains the most fragile part of our runtime architecture. If we were starting over, we would stress-test the loading pipeline with worst-case dungeon configurations before building anything else on top of it.

The Generation Layer: Custom Code

The dungeon generation engine, the level grammar, the enemy ecology system, and the seed architecture are all custom code, written in GDScript with performance-critical sections as GDExtension C++ modules. We considered several existing procedural generation libraries early in development and ultimately decided not to use them for the core pipeline.

The reason is architectural control. Procedural generation libraries typically make decisions about graph representation, constraint evaluation order, and backtracking strategy that are baked into their design. Our grammar needs specific evaluation behavior: constraint priorities we can tune, a deterministic evaluation order compatible with our seed system, and a backtracking model that logs constraint relaxations for analysis. Adapting an existing library to those requirements would have been more work than writing the system directly.

This is not a recommendation for other studios. For teams with less specific requirements, a well-maintained library would likely be faster to ship with. Our requirements were unusual enough that custom code made sense for us. The cost is roughly six months of engineering time that we could not spend on other features.

The Behavior Tree Runtime

We use a modified version of an open-source Godot behavior tree plugin as the base for our enemy AI runtime. The plugin handles the tick loop, the node evaluation order, and the blackboard (per-agent memory store). We extended it with our trait-grafting system: the ability to dynamically attach and detach subtrees from a tree structure at runtime, which is what the behavior assembly pass requires.

The modification was non-trivial. The original plugin assumes a static tree structure initialized at scene load. Our assembly pass attaches trait subtrees after initialization, which required changes to the plugin's internal reference management and to blackboard scope handling for attached subtrees. We contributed some changes back to the plugin's repository; the most specialized modifications were not general enough to be useful to other projects.

The Data Pipeline: Direct Configuration

The grammar rules, the enemy role taxonomy, the behavior trait library, and the ecology composition weights are defined in JSON configuration files that we edit directly. There is no visual rule editor. Priya and I write grammar rule changes in a text editor, check them in, run the generation test suite, and observe the output.

This is probably the most glaring tooling gap in our current workflow. A visual constraint editor with immediate generation preview would reduce the cycle time for grammar changes significantly. When Priya wants to test a new room composition rule, the current loop is: write rule, run test suite (roughly four minutes), inspect fifty sample outputs, identify issues, revise, repeat. A visual tool with instant preview could compress each cycle to under a minute.

We have not built that tool yet because it is a meaningful engineering investment that would take Omar away from the generation engine during a period when the engine itself still needs attention. This is a real prioritization tradeoff. The tooling gap slows Priya's iteration speed and that compounds over time as a design cost we are absorbing.

Generation Monitoring

We run a generation test suite that produces five hundred dungeon instances per test configuration and measures structural metrics: room role distribution, critical path length variance, constraint relaxation frequency, encounter threat value distribution by depth tier, and layout uniqueness scores. The suite runs in about eight minutes on a development machine and is part of our standard build validation step.

The monitoring output is console-based: plain text log files that we grep and aggregate with scripts. This is adequate for current scale but will become inadequate when the grammar's complexity expands further. At some point we need a proper metrics visualization layer. That is another tooling investment in the queue.

What We Would Change

With the knowledge we have now, we would invest more time in the loading pipeline before building on top of it. The resource streaming fragility has cost more debugging hours than a more extended upfront investment would have justified.

We would build the grammar visual editor earlier. The compounding cycle time cost of text-based rule editing has been larger than we expected at the start. A small design tool investment early would have returned value by now in Priya's iteration velocity.

We would keep the custom generation engine. This has been the right architectural call for our specific requirements, and the control it gives us is worth the six months of engineering investment. The behavior tree base was the right call too: using and modifying an existing plugin rather than writing from scratch saved meaningful time without locking us into a structure we could not modify.

We cannot tell you whether our toolchain is right for your studio. It is calibrated to our specific constraints: a three-person bootstrapped team, an unusually specific generation architecture, and a production cadence that prioritizes system quality over feature velocity. Your constraints will be different and the right tools will reflect that.