← all_posts
Engineering2026-09-25·19 min

Visual Novel Mode: Giving a Language Model a Stage

RoleCall's VN mode turns streaming roleplay into sprites, backgrounds, choices and endings. The hard part was never the renderer. It was deciding what happens when the model, the author and the player disagree.

A scene mock from The Understudy, a test VN: two sprites in a dressing room with a dialogue box

In May we shipped visual novels on RoleCall. You build a cast with sprite sets, a list of locations with backgrounds, a story graph with choices and endings. A player gets a stage: a nameplate, a lit sprite, a dialogue box, music. Behind it is the same roleplay engine as everything else on the platform, which means the lines on that stage are written live by a language model.

That combination is the whole point and the whole problem. A traditional VN is a script with a renderer attached. Ours is a renderer attached to something that improvises.

Nemo and Chi built the first runtime, the stage tools and the bundle format in the spring. In late August I picked it up, found that a lot of what the editor let you author was never read at runtime, and spent the next three weeks making it actually work. VN mode is now about 105,000 lines of TypeScript, a third of it tests, and most of the commits since August are mine. Almost none of them were about drawing things on screen. They were about three parties who each think they own the scene: the author who wrote the story, the model performing it, and the player who can type anything at any time.

This post is about how we settled those disputes.

One decision that everything else follows from

A VN on RoleCall is a character that carries a vn_data extension. There's no per-chat toggle. A VN character is always played as a VN and a plain chat is never upgraded into one.

And one chat is one playthrough. This is the load-bearing decision in the spec, and most of the design falls out of it:

  • The transcript is the save file. There's no save/load system. Beat snapshots plus our existing branching ("takes") already give scrubbing and forking, and a slot system would be a second, worse mechanism for the same thing.
  • Replay means a new chat. A finished run stays readable forever.
  • Rewind stops being a convenience. It's the only way a player recovers from a choice they regret, so it has to be exact. That forced us to snapshot evaluated state at every beat instead of recomputing it on restore, because recomputing would re-fire objective triggers and re-swap backgrounds every time you step back.

The other founding decision came from the launch notes: VN mode is an overlay on top of normal chat, not a separate mode. The player can always type. That sounds friendly. It's also the source of most of the bugs below.

How prose becomes a stage

During play the model writes normal prose with stage directions inline:

[VN_ACTION]{"type":"SET_BACKGROUND","location_id":"dressing_room","variant":"night"}[/VN_ACTION]
[VN_ACTION]{"type":"FOCUS_CHARACTER","character_id":"vespera"}[/VN_ACTION]
[VN_ACTION]{"type":"SET_EXPRESSION","character_id":"vespera","expression":"smirking"}[/VN_ACTION]
Vespera: You wore my costume. You stood in my light and you were good in it.

There are 24 action types: backgrounds, music, sound effects, expressions, sprite movement across seven stage positions, focus, choices, state changes, items, objectives, minigames, map navigation, characters entering and leaving. The parser pulls the tags out of the stream, strips them from what the player reads, and validates every action against what actually exists in the VN: a background the author never drew, a character not in the cast or a music track that isn't there is dropped rather than rendered as a broken image.

Why the default is tags, not tool calls

The obvious modern design is native tool calling: give the model a set_expression tool and let the provider handle the structure. We support that, and a player can pick it. It's not the default, and the reason is positional.

The reader advances one beat at a time and picks a nameplate and a lit sprite per beat, so beat boundaries have to be locatable inside the prose. Tagged output carries them: the model writes the tag in position, in the same pass as the text around it. Native tool calls can't. In our tool pipeline every tool runs to completion and the visible prose comes from a separate final request, so a FOCUS_CHARACTER call has no position inside text that didn't exist when it was made. Tool-call scenes got one spotlight for the whole turn, and narration kept landing under the previous speaker's nameplate.

So the transport is a user choice with no automatic mode, and tagged is the fallback for missing or malformed settings. That's a correctness call, not a preference. A silent automatic switch between the two would have made the failure invisible.

"The model isn't emitting tags"

This was the first bug report and the one I learned the most from, because it was never one bug. Over two days it turned out to be five:

  • A leftover test author's note sat at depth 0, the very bottom of the prompt and the most influential position there is.
  • The worker was running stale prompt code. Prompt assembly doesn't hot-reload, so some of my fixes weren't being tested at all.
  • Tags were stripped before the reply was saved. So in the chat history, the model only ever saw itself writing plain prose. The history is part of the prompt, and it was quietly demonstrating the opposite of what the contract asked for.
  • SHOW_CHOICES was named in the contract but its shape was never shown. The model invented a plausible schema, and the parser rejected it.
  • "Put FOCUS_CHARACTER on its own line" was taken literally, and written as prose.

The fixes were to keep the raw emission and show it in the story log, to put actual JSON shapes in the contract, and to phrase instructions positively. We'd added a prohibition along the way, emissions dropped, and we reverted it. I'd call "positive phrasing works better" a hypothesis rather than a finding, but it's cheap to apply.

Placement mattered even more than wording. On one fast model, putting the contract before the chat history produced zero stage actions, and putting it at the very bottom produced actions plus the model reciting its instructions into the dialogue box. Then a real, heavily built preset put the VN block at message 39 of 61, behind about 40,000 tokens of other instructions. We added a macro so preset authors can place the contract themselves, and on the same model that went from 0 actions in 3 turns to 10 actions in 6.

The most expensive lesson came from a character who wouldn't leave the stage. We reworded the exit instruction four times before dumping the full prompt, and found the list of who was on stage had never been in it: for the first seconds after page load the stage was empty, so the prompt said nobody was there. Dump the prompt before you reword it.

Some fixes were just meeting the model halfway. Cast ids were timestamps, and models mistyped 13-digit numbers constantly, so any character_id field now also accepts a display name. Models used CLEAR_FOCUS to mean "leave", so the contract had to separate "holds the stage" from "takes them off the stage". And NAVIGATE_TO and ADVANCE kept getting confused until we noticed both tool descriptions began with the word "Move".

Models forget to stage

The next thing live testing showed us: models routinely write a labelled line (Vespera: ...) without issuing any staging command at all. The speaker was talking from an empty stage. The fix is unglamorous. If a line is labelled with a cast member who isn't on stage, their sprite stages itself with its default expression. The runtime doesn't wait for the model to remember.

That only works because beat detection doesn't depend on the model either. Every paragraph is its own beat. A Name: label wins, then an explicit focus action, then nobody. A paragraph that opens with another cast member's name, or has no quote marks, doesn't mention the speaker and has no first-person pronoun, drops out of the speaker's beat into narration. These rules are deliberately model-independent, because we'd measured a model that emits no stage commands at all and the stage still had to read correctly.

Timing was its own bug class. Stage changes used to fire as soon as the stream finished, so the reader would see a background swap or a character exit several beats before the text that caused it. Two expressions in one reply both applied to every beat, so the sprite always settled on the last one. Expressions are now anchored to the paragraph they were written beside, and exits, transitions and background changes wait until the reader actually pages to their beat. The rule is that anything the reader sees waits for its beat, and bookkeeping (state changes, items) applies immediately, or the game state would depend on how far someone had scrolled.

Expressions have one more failure mode. A model writing choices will invent expression names, because it doesn't have the sprite set in front of it. An unknown name falls back to the character's default face rather than rendering nothing. The protagonist gets a longer chain (the face the hovered option asks for, then the face the model set, then a "considering" face, then the default), because a protagonist blanking mid-scene reads as a broken image rather than a missing option.

The story graph that couldn't move

Authored VNs have a story graph: scene nodes, choice nodes, condition nodes, free-play nodes and endings. For months the graph reached the runtime but never reached the model, and when I finally traced how a playthrough moves through it, I found three facts that were true at the same time:

  1. SHOW_CHOICES was the only action in the entire union that touched the graph. There was no "advance", no "go to".
  2. The function that moves the graph had zero callers outside the store.
  3. Inside the store, only the choice handler called it.

So the complete path for a story to progress was: model emits a choice, player clicks, graph advances. On a model that emits no VN actions (and we measured one that emits none) a playthrough never left the first node for the entire session. It also produced a double prompt in the most common authoring shape: to leave a scene whose next node is a choice, the model had to invent button text for a choice whose only job was to reach the node that then showed the author's real options.

Two additions fixed it:

  • ADVANCE. No argument in the common case. It follows the current node's exit, and if that lands on a choice node the runtime posts the author's own options. The model's job dropped from "construct a valid choice array with correct node ids" to "this beat is done".
  • A player-facing Continue. Shown when the current node has exactly one wired exit and nothing is pending. This takes model compliance off the critical path entirely. The graph can always move, because the player can always move it.

The graph also goes into the prompt now, and what we left out of it mattered as much as what went in. The model sees only the node the player is on and the exits it may legally take, never the whole graph, because that block rides on every VN turn. It doesn't see branch conditions, since the evaluator owns those and the model would only narrate the mechanic. With a single exit, the prompt asks for a bare ADVANCE and doesn't mention an id at all, because naming an id the model doesn't need to supply is an invitation to invent one.

Who owns the beat

Each scene has a mode that decides who owns it: the author, the model, or both.

ai_drivenhybridscripted
Authored linesignoredplay firstown the beat
Model's textowns the boxtakes over afterhidden while lines are unread
Player inputlivelivelocked while lines are unread
Model's stage actionsappliedapplieddropped while lines are unread

The question that separates them isn't "is there authored text". It's can the player interrupt it. Hybrid is an opening the model continues from. Scripted is a cutscene.

The first version of scripted mode gave authored lines an unconditional hold on the dialogue box and stopped there. The player could still type. So the message was sent, a generation ran and was billed, the reply was written to the chat and the story log, and then it was thrown away unseen because the scripted queue kept the box. Its stage actions still fired, so on a capable model the background and expressions changed with no text explaining why.

The lesson generalises: making authored content win is only half a decision. The other half is what happens to the input that content is now refusing. The rule now lives in one shared predicate that three far-apart consumers ask (the overlay, the composer, the action processor), because a disagreement between any two of them is invisible until a player hits it.

Authored lines also used to exist only on the stage. They were missing from the transcript and from the model's context, which is why a hybrid scene could ask the model to "continue from where the authored lines stop" while showing it none of them. They're now replayed as synthetic assistant messages, following the same pattern as a character's greeting: derived from state, stable id, never written to the database, merged into the transcript by timestamp.

246 fields, 18 of them ignored

In August I swept every field in the VN content model against the code that actually plays a VN. 246 fields, 18 with zero runtime readers. An author could fill them in, they would round-trip through export and import perfectly, and they would never affect a playthrough.

Every one failed the same way: authored data reached a boundary that dropped it, with no error and nothing in the creator saying so. A mapping function that listed six fields and discarded the rest. A render filter. An action switch that never received the action. Nothing in the stack treated "the author filled this in" as a fact worth checking.

The field the automated sweep missed is the instructive one. Scene content_type, the mode from the section above, showed runtime hits and passed the filter, because the name collides with unrelated fields elsewhere in the codebase. The grep narrows the search; it doesn't settle it.

The spec that came out of the sweep resolved every one of the 18: built, repurposed, or deleted. Some deletions are worth explaining:

  • Choice timers. Timers are a tension mechanic, and tension doesn't survive a fork where the surrounding beats take seconds to generate.
  • play_count and completion_count on the VN. The VN lives as a JSON blob on the creator's character row, saved wholesale. A counter inside it means every player who finishes a VN writes to content they don't own, through a route that replaces the whole blob. That's write contention and a permissions problem wearing a counter's clothes. Both are queries now.
  • Inventory capacity. Enforcing it needs an overflow UX nobody asked for, and a model-driven GIVE_ITEM that silently fails is worse than an unbounded bag.

A spec that only adds is how the schema got there in the first place.

Choices, endings, and not lying to the player

Once the graph could move, choices had to mean something. Choice nodes now lock the composer, since a choice point is a decision and until then it was a suggestion you could type straight past.

That created a softlock. Options whose conditions fail were filtered out, so with every option filtered and the composer locked, there was no way forward. The fix was to show locked options disabled instead of hiding them, which also turned out to be the better design: a player who never sees the locked option never learns the fork existed, and that's most of what makes relationship stats feel like they matter.

What a locked option says took more care than I expected. The easy move was to reuse the editor's condition-to-text renderer. That's the wrong source, because a condition's operands are identifiers the author chose for the machine. Worse, some conditions are judged by a model from an author-written prompt, and those prompts routinely state the author's opinion of the player: "the player has been cruel to the stagehands." Shown on a disabled button, that breaks the fiction and reads as an accusation. So locked options get a terse badge (Trust 40) only for numeric requirements at the top level of the condition, an optional author-written line, or nothing at all. Under or a badge would claim a requirement that isn't required. Under not it would state the inverse of the truth.

Endings got the same honesty treatment:

  • No quality tier. An early draft had bad, normal and true endings. Route-based VNs don't rank their endings, and "true end" carries a specific genre meaning (locked behind seeing the others) that we don't build. Authors get one flag, canonical, which is display only.
  • An honest count. The endings list says "3 of 7" without saying what's missing, and there's no secret-ending flag to hide one from the count.
  • Retired endings. If an author deletes an ending a player already reached, the row stays, marked as no longer part of the story and excluded from the count. Dropping it would erase a finished run because somebody edited content afterwards.

Meters with a rubric

Relationship meters started as a name and a number. Now a meter carries a declaration: what it means, what earns a rise, what costs, and how far one ordinary move may shift it. Trimmed, from a court-romance VN I built to stress the engine:

{
  "name": "affection",
  "meaning": "How much in love this character is with the player",
  "positive_prompt": "honesty, especially unflattering honesty; keeping a promise exactly as stated; choosing him and saying so.",
  "negative_prompt": "evasion; charm answered with charm; broken or softened promises; choosing another and letting him find out secondhand.",
  "min_change": 2,
  "max_change": 15
}

The bounds are the part that makes it playable. Without them one good line can swing a meter from cold to devoted, and every gate the author built above it stops mattering.

The editor: every edge is a port

The original story editor was a flat list where links were set by picking node ids from a dropdown. Blind, and unusable past a dozen nodes. Its conditions rendered as read-only JSON. Validation was a button you had to remember to press, and reachability was checked as "not referenced by anything" rather than as a traversal, so disconnected islands that referenced each other passed.

The replacement, Flow, is a node canvas. The data model has four kinds of edge field, and each becomes exactly one visible port on a node card, so wiring is a drag. A condition renders as a sentence where every token is a picker (IF relationship Nadia trust >= 40), with variables drawn from the story's starting state, which is the only mechanism that keeps the graph and the state in sync. Validation is continuous and calls the same function the publish check uses, so the canvas and the publish button can't disagree about whether a story is playable.

Nobody had ever positioned a node, so the first open has to lay the graph out. A BFS layering (column is depth from START, row is index within that depth, cycles tolerated by visiting each node once) is about sixty lines and saved us adding a layout dependency. Unreachable nodes get parked in a gutter below the graph, so "floating down there" reads as the problem it is.

Testing a story you didn't fully write

An authored VN is still half improvised, so "does this route work" can't be answered by reading the graph. Test mode, built in September, plays the story exactly as a reader would, in an iframe of the real scene page inside the editor. It's an iframe and not an embedded component because Tailwind breakpoints read the window, not a container, so a player drawn in an editor column could never show what a reader on a phone sees. The frame gets real viewport presets instead.

Alongside it, the run keeps a trace of every node it landed on and which exit it took to get there, including exits the model picked by name. When a test session reaches an ending along a route the story declares, that route is marked tested automatically. Force-exits (the author's hand on the wheel) are reported but don't count. Editing a node's prompt doesn't untick the routes through it; only a structural change does.

We planned auto-play, with a model driving the player side to walk routes unattended. We dropped it. It costs model calls, and Force exit plus the reader's own Continue already cover walking a node by hand.

Assets: the boring part that ate a week

A VN is a lot of art. The .vnpack bundle format exists so a story can be built offline (write JSON, drop in images) and imported whole: a zip with vn.json, a manifest, and content-addressed assets, so identical bytes referenced from several places are stored once. On import every image is re-encoded server-side, which strips metadata and defeats polyglot files, and audio goes through an extension allowlist.

I dogfooded the format hard. The court-romance VN went through 39 pack versions in about two weeks and ended at 22 cast members, 43 locations and 276 expression sprites. Early on, a 126-asset pack exported at 256 MB and our own importer refused it. The cause was the upload route, which had been re-encoding all VN art to PNG. Switching it to WebP took a 2752×1536 background from 8.2 MB to 310 KB, a sprite with alpha from 4.7 MB to 237 KB, and the whole pack from 256 MB to 13.6 MB. We left the size limits exactly where they were. That one change did more for load times than anything in the renderer.

Sprites that stay the same person

The art pipeline went through two generations, and the second one taught me more.

The first was local: an SDXL anime checkpoint in ComfyUI with a style LoRA and a face-detailing pass, a fixed seed per character, six emotions each. Background removal used an anime-tuned segmentation model, keeping only the largest connected region, which on its own removed floating artifacts from half the sprites. The recurring failure was that every noun placed near the head in a prompt became a literal object there: a sun, an extra pair of ears, once an eyeball.

The second generation used a multimodal image model with reference images, and consistency became the whole game. With one reference (the portrait), costumes drifted from frame to frame. With two (the portrait plus the character's own neutral sprite as a costume anchor) it was decisively more consistent, with a catch I wrote down at the time: the anchor locks in whatever the anchor got wrong. A warm tint, a stray forge in the background, once a faint second head. The fix was an automated gate that rejects a frame whose corners aren't clean white and rerolls it, up to three times. An eight-expression set came out to about $1.10.

The last surprise was scale. Normalising sprites by bounding-box height made heads grow on tighter crops, so a character visibly changed size when their expression changed. Normalising by head width instead took the spread in head size across one character's set from 58% to 1.4%.

A main character ends up with up to 28 expressions, a minor one with 8. Each expression carries a short note in vn.json saying when that face is used ("wounded pride, betrayal, rejection"), and that note is what the model reads when it picks one. The prompt lists expressions only for characters currently on stage. With 17 cast members it used to be 17 lines, and the cast lore block shrank from 17 character sheets to the one or two who are actually present.

The most recent test story, The Understudy, is a gothic romance set in a theatre: seven cast members, eight locations, 31 scenes, 44 story nodes, five routes and seven endings. Its season is eight nights long with five suitors. You can't see everyone, and that's the point. Its final choice gates on flags rather than meters, so a route you actually played is always available at the end, whatever the model did to the numbers along the way.

What I'd tell someone building this

Give the player a way to move the story that doesn't depend on the model. Every mechanism that required model compliance to progress eventually stranded a run on some model. Continue buttons are boring and they're the reason runs finish.

Positional information decides your transport. If the UI needs to know where in the text something happened, inline tags beat tool calls, whatever the tool-calling benchmarks say.

Audit the fields, not the features. The 18 dead fields weren't bugs anyone reported. They were promises the editor made that the runtime never kept, and the only way to find them was to check every field.

When authored content wins, decide what happens to what it's refusing. Every mode, lock and gate in VN mode has a second half, and the bugs lived there.

The last VN beta flags came off on September 8th, so the canvas and the graph-aware prompt are now simply what every VN author gets. If you want to see what a model does when it's handed a stage, go make one.