← all_posts
Engineering2026-09-25·12 min

RoleCall, Six Months Later: From Invite Codes to Paying Customers

What three people shipped between April and September: a public launch with paid plans, characters with minds, undo for an entire world, a tabletop mode, visual novels, and one week where our own retry logic tried to take down the database.

In April I wrote a run of posts about how RoleCall works inside: the multi-agent turn, the Story Director that plans behind your back, the Compendium, the immersion trackers, and inference routing. Those posts describe an architecture. This one is about what happened when we pointed that architecture at real users and kept building.

Some numbers for scale. On April 12th Chi set up a new monorepo and ported RoleCall into it. Since mid-April that repo has taken about 4,500 commits from three founders (Chi, Nemo and me) plus one contributor, and 623 database migrations. The app version went from 2.178 to 2.717. By commit type, fixes outnumber features by about 5 to 3.

The shape of the company changed first

RoleCall used to be one app. It's now several apps on one shared Supabase database:

  • RoleCall is the product: scenes, the editors, the whole stage.
  • PlotLight is public discovery for characters, presets and visual novels, plus PlotPoints (more on that below).
  • Stage Whispers is a social feed that's still taking shape.

The split isn't aesthetic. Card payment processors are cautious about a checkout form sitting next to user-generated adult content, so the paid product and the public discovery surface live in separate apps with separate domains. Keeping them on one database means a character published on PlotLight is the same row you play in RoleCall. We didn't invent this pattern. Other creative platforms with user uploads split the same way, for the same reason.

Opening the doors

At the end of June RoleCall went from invite-only to open signups, and paid plans went live with it. I did a lot of the billing work, and the design choices there were more interesting than I expected billing to be:

  • Plans grant weekly token pools, not a monthly dollar balance. Each tier funds a ladder of pools that refill every Sunday at 00:00 UTC. Higher pools cascade down to cover cheaper models, and the top tiers include the lower pools natively so nobody spends premium tokens on a cheap model.
  • The entry tier counts requests, not tokens. A request-metered plan with its own tab in the model picker is much easier to explain than a token budget to someone who has never heard the word "token".
  • The free tier keeps the whole app. Every feature works on free, with a set of free models and a monthly request allowance. Gating features would have split the product into two products to maintain.
  • The balance card has to add up. Top-ups, courtesy grants and gift tokens are listed separately, the refill time is shown in your own timezone, and "subscription renews" is labelled so it can't be confused with the weekly refill.

The token pool is charged inside the narrator engine itself, per generation, which is also where the usage ledger records what each call actually cost us. So pricing decisions can be made from what calls really cost rather than from estimates.

Understudies: beta testing as a toggle

Once there were paying customers, we couldn't keep shipping half-verified features straight to everyone. We'd been in an "exit freeze" before launch, and coming out of it we introduced Understudies: a single toggle in Settings that opts you into beta features. No invite, no waitlist, flip it off any time.

The rules for engineers are strict and short:

  • Every brand-new user-facing feature is beta by default, unless it's a bug fix, a copy tweak, or a change to a surface that's already live and verified.
  • A beta feature gets a key in one registry and is checked at the UI and at every server surface. UI-only gating is not gating.
  • Checks fail closed.
  • Killing a broken feature is flipping its key to false. Graduating it is deleting the key and every call site.

It's a boring mechanism and that's why it works. TTRPG mode, preset choice groups, the regex builder and a dozen other things went through it, and on September 19th every remaining key graduated in one pass.

Characters with minds

The biggest system I built this cycle is Character Minds. The Compendium remembers facts about the world. Minds is about what a character remembers and how they feel about it, which is a different thing.

  • Memories per witness. A scene with three characters in it produces three memories, each with that witness's own emotional reading. Subjective memory falls out for free.
  • Mood on five axes: valence, arousal, dominance, anxiety and focus. Each drifts toward the character's baseline as story time passes.
  • Memories recolour. A memory is re-stained by the mood it's recalled in. A happy memory of someone becomes a sad one after they die, permanently.
  • Drives. Wants and fears are tracked explicitly, and resolve on their own when the scene fulfils, thwarts or abandons them. A dream that's lost leaves a scar that shifts the character's resting state.
  • Theory of mind. Characters track how they see each other (affinity, trust, the dynamic between them) separately from what's actually true. Your tsundere can be wrong about how much someone likes them, and the narrator will play it that way.
  • Twelve personality dials govern how all of the above behaves. The same scene history produces a different inner life for a ruminator than for a forgetter.

Minds shipped alongside a semantic layer I built the same week: pgvector with an HNSW index over one polymorphic embeddings table, so characters, lorebooks, presets, personas and the rest of the library share one retrieval path under the same row-level security as everything else. Character memories carry their own vectors. The same layer gave PlotLight semantic search and RoleCall semantic lorebook retrieval.

Minds also feeds a Scene Frame, a compact per-turn card of where and when the scene is, who's present, and what each person there currently wants and feels. It's the cheapest way we've found to keep a model oriented in a long scene: not more history, a better summary of the present.

Undo for an entire world

RoleCall turns write to a lot of state: lorebook and Compendium entries, Story Director plans, storyboards, the map, 22 immersion trackers, character minds, dice rolls and shop purchases. For a long time, deleting a message deleted the text and left all of that behind, so the world kept remembering things that "never happened".

In June we shipped turn-wide undo: every write a turn makes is journaled, and deleting a message, a branch, a tail of messages or a single swipe rewinds the world to exactly where it was before that turn. A few rules made it trustworthy:

  • A later turn's edits win. If a newer turn you're keeping changed something the deleted turn wrote, the newer value survives.
  • Bulk deletes coalesce. A hundred-message rewind restores each entry once, not a hundred times.
  • Swipes are isolated. Deleting one swipe rewinds only that swipe.
  • Rollback never blocks the delete. If the rewind hits a problem, the message still goes, and the rewind reports what it couldn't do.

It's one of those features users never think about until it's missing, and then it's the only thing they think about.

More ways to steer

A lot of the cycle went into giving people control over the story without breaking character:

  • Orison, the stage assistant Chi and I built. You open it with /btw from anywhere and it acts on the surface you're on: clicks panels, reads earlier turns, hands a rewrite to post-production, plants arcs for the Story Director to pay off later.
  • The Interpreter, Chi and Nemo's per-turn director. It never writes fiction. It reads your latest move and tells the writer the shape of the next reply: compress travel, dwell in a moment, keep the scene from snapping shut. It also treats a rejected swipe as a hard signal about what not to do next.
  • The Dramaturg, which I built in August. It's a second assistant whose subject is the fiction rather than the app: "why did he answer like that", "is this consistent with her card". It's strictly read-only, and it answers from the material the scene actually ran on. To make it useful I gave every in-app agent a scene-search tool, so asking about something fifty messages back jumps straight to it instead of reading forward from the start.
  • A preset macro engine with state hooks, in-reply captures, model-decision macros, a lint pass, and a zero-cost playground where preset authors can run their macros against a fake world without spending a token.

Two new ways to play

TTRPG mode is Chi's: server-backed character sheets for D&D 5e, Pathfinder, Draw Steel, a Disco Elysium-style system or one you build yourself; dice that resolve server-side; conditions the engine enforces so a "+2 blessed" can't drift; shops with quote-then-confirm purchases. It went through Understudies before graduating.

Visual novels are mine. You author a cast with sprites, locations, a story graph and endings, and the model performs it on a stage. That one got its own post, because the interesting part is what happens when the author, the model and the player disagree.

PlotPoints: the arena found a home

The RP benchmark arena from my benchmark post started life on its own subdomain. Chi built PlotPoints, a permanent home for it inside PlotLight, and on April 30th I moved the live multi-turn round over mid-vote, importing the 507 votes it already had. That round closed in June at 1,943 votes from 482 voters, on top of 2,013 for the single-turn round before it.

The most interesting result from round two was an inversion. Across the 11 models voted on in both rounds, the multi-turn ranking was negatively correlated with the single-turn one (Spearman ρ = −0.24). Gemma 4 26B, the surprise #1 in my benchmark post, finished the multi-turn round at #15. GPT-4.1, dead last with single-message voters, climbed into the top five once people read whole sessions. Single-message votes measure charm; twelve-turn votes measure stamina.

Round three is running now with a refreshed 21-model pool and a separate 40-model After Dark track, because the data had already shown that adult scenes reshuffle the leaderboard. Its judge-scored results so far have one finding that matters for the end of this post: the seven RP-specialist fine-tunes in the After Dark pool, the models marketed for exactly this, ranked #34 to #40 out of 40. Human votes on that track may yet disagree.

The week our retries attacked us

At the end of August our database started logging millions of errors a week and crash-looping daily. It took more than a week to fully understand, and nearly all of it was self-inflicted, which is the most useful kind of incident to write up.

The pattern: at the end of a turn, the client captures the world state and asks the database to "seal" it. The seal re-derives the same values server-side and compares. If they differ, someone else wrote in between, so it rejects with a serialization error and the client retries. That's optimistic concurrency, and it's fine as long as a mismatch means "try again" rather than "this will never match".

It stopped being fine because several mismatches were permanent:

  • The client retried 409s. A conflict isn't transient, so retrying one just repeats it. That alone was an 8x amplification. Now only 5xx retries.
  • The comparison included server-set timestamps, which by definition never matched the client's copy.
  • A revision trigger fired on fields the turn itself writes, so the seal invalidated itself.
  • Retries re-committed intents prepared hours earlier. One prepare was followed by over 24,000 commit attempts. Commits against intents older than 60 seconds are now rejected. For scale, a healthy commit lands within about 8 seconds.
  • Nothing bounded the retries. A per-chat circuit breaker now opens after five consecutive races.

And the big one: the seal read the whole ledger in one query, and the API layer silently capped responses at 1,000 rows. Any chat past 1,000 ledger rows computed its cutoff from a truncated list, which made it permanently wrong, which made the seal fail forever. No error, no warning, just a quietly short array. The fix was pagination. The lesson is to assume any list you didn't paginate is lying to you.

The last piece took the longest to find, because the retries weren't coming from our code at all. Our conflicts were raised with Postgres's serialization-failure code, the one that means "this is transient, try again". The API layer in front of the database believed it, and quietly retried them itself, indefinitely. Re-raising application conflicts as a plain HTTP 409 stopped it within minutes. Never use a retryable error code for a conflict that will never resolve, because something in your stack will take you at your word.

One storm we resolved by removal rather than by guard. We took a class of consistency checks on summaries out entirely and made them last-write-wins, invalidated when history genuinely changes. A check that can't be satisfied isn't protecting anything. And while we were in there, we found the write ledger was re-snapshotting a full embedding vector on every write. Storing it once took a turn's ledger payload from 753 kB to 23 kB.

Three more things from the write-up I'd tape to any monitor:

  • Your query stats may not count failures. Ours only recorded successful statements, so the worst queries were invisible in the tool we reached for first.
  • Verify what's deployed, not what's on the branch. A correct fix that isn't serving yet looks exactly like a wrong one.
  • Reconcile before you panic. We checked every stuck write afterwards and every one had in fact committed. No data was lost.

Measure the wait before you fix it

In September I went after the time between pressing send and anything happening. The obvious suspect was the typing indicator. It wasn't. It drew in about 85 milliseconds. When I actually traced a send, the time went to three things nobody had looked at: a settings write that ran on every page load and cost about two seconds, the durable write of your message (one to two and a half seconds, growing with chat length), and, on long chats, a full reload of the history on every send, because a cache's 30-second idle window expired inside a normal three-minute turn.

The fixes were small once they were visible. The message write and prompt assembly didn't need each other, but ran one after the other: on a 204-message chat, 5.3 seconds passed before the model was asked for anything. They overlap now, and the same chat takes 2.5.

The cache was subtler. Its 30-second window wasn't arbitrary: holding a second copy of a long chat in memory was a top cause of the browser tab crashing on Android phones. So the fix wasn't just a longer timer. The copy is now dropped the moment the tab is hidden, which is exactly when a phone reclaims memory, and the idle window is long enough to survive a slow turn. None of it had anything to do with the model.

What's next

Our own inference is next: the first package for it landed in September. I've had one taste of what that involves already. Around the start of June I spent eight days trying to make Strovolos, the host of our Discord, run on my own GPU instead of an API, and that story is its own post. It and the PlotPoints result above point the same way: a fine-tune moves a model toward its data, it doesn't lift it above its base.

The thing I'd underline from these six months: almost none of the hard problems were about the model. They were about state. Who owns it, when it's allowed to change, how to undo it, and what to do when two parts of the system disagree about it. The model is the most visible part of RoleCall and the least of our problems.