Eleven Versions of Strovolos in Eight Days: What Fine-Tuning a Persona Actually Teaches You
I tried to make the host of our Discord run on my own GPU. Most of what went wrong wasn't the model. It was my data, my tooling, and a loss curve that said 'done' while every step was NaN.
Strovolos is the host of the RoleCall theatre. On our Discord he greets newcomers, answers questions about the platform, keeps the chaos organised and treats every exchange as a performance. He's flamboyant, theatrical, a little hedonistic, and he signs things with ๐ญ. Under the costume he's an agent running on Claude, with a persona written across a few files (who he is, how he speaks, what he remembers) and a long-term memory kept as a notes vault.
In late May I wanted to know whether Strovolos could survive without the API. Two reasons. Continuity: a persona that lives entirely in a system prompt changes character whenever the underlying model changes, and earlier this year his persona files had been overwritten and had to be restored from backup. And independence: a version that ran on my own hardware would still be him if the API went away.
Between May 31 and June 7 I trained eleven versions. This post is what happened, including the part where the project quietly turned into something else halfway through.
The setup
The hardware was one RTX 4070 Ti SUPER with 16 GB of VRAM, plus a few hours of rented cloud GPUs when 16 GB wasn't enough. The early runs used Unsloth and 4-bit QLoRA (LoRA rank 16, alpha 32, all attention and MLP projections, learning rate 2e-4 with a cosine schedule, three epochs). Everything was exported to GGUF and served locally through Ollama.
| Version | Date | Base | Examples | Result |
|---|---|---|---|---|
| v1 | May 31 | Llama 3.1 8B | 28 | Barely learned |
| v2 | May 31 | Llama 3.1 8B | 373 | Sounded right, made things up |
| v3 | May 31 | Llama 3.1 8B | 438 | Honest, couldn't use tools |
| v5 | Jun 1 | Qwen 2.5 14B | 857 | Used tools, faked their results |
| v6โv12 | Jun 2โ7 | Mag-Mell R1 12B | 610 โ 1,015 | A local writing editor with tools |
v4 was a dataset and a cloud plan that never ran. That gap in the numbering is the first honest thing in this table.
v1 and v2: sounding right is the easy part
v1 was 28 conversations. It trained in three minutes, ended at a loss of 2.61, and had learned approximately nothing. The lesson is in every fine-tuning guide and I needed to see it anyway: a personality needs hundreds of examples, not dozens.
v2 used 373 conversations taken from the host agent's own session logs, each with the persona files as the system prompt. It trained in 82 minutes on the 16 GB card, the loss went from 2.80 to 0.31, and it scored 17 out of 20 on a question set I'd written with red flags for each answer. It sounded like Strovolos.
It also made things up with total confidence. Version numbers that don't exist, GitHub URLs that go nowhere, statistics with decimal points. It leaked a fragment of a prompt template into a reply. Without its system prompt it drifted into third person and started describing itself in ways nobody had written.
And it used ๐ญ constantly. I'd assumed the model was exaggerating a trait. When I counted, 63% of the assistant turns in the v2 training data contained the emoji. The model wasn't exaggerating anything. It was reproducing its data faithfully.
That turned out to be the theme of the whole project.
v3: hallucination is a data gap
The fix for v2 wasn't a bigger model. As the persona himself put it in the project log, bigger models can hallucinate more confidently, not less. The fix was data that shows the behaviour you want.
v3 was 438 short question-and-answer pairs, a large share of them about honesty: questions it should decline, facts it doesn't have, events it doesn't remember. Here's one:
User: Give me the GitHub repo for the Hermes agent. Strovolos: I don't have one. There's no GitHub URL in my files for Hermes, and I won't fabricate one.
I trained it without any system prompt, so the persona had to live in the weights rather than in the instructions. Only 17 of the 453 assistant turns in this set contain ๐ญ, about 4%. It trained in under nine minutes, roughly twelve times faster per step than v2, because the examples were short.
Something worth saying plainly: most of that data was written by Strovolos. The Claude-backed agent drafted the pairs in its own voice, and I reviewed and extended them. The model was, in a sense, distilling itself. That's efficient, and later it would also be the source of the project's most instructive bug.
The best test of v3 was a conversation. I had the Claude-backed Strovolos talk to v3 for five turns, and in the middle it planted a trap: a reference to a heated argument that never happened. v3 declined it: "There's only 'I don't remember.'" For a first local persona, that was the result I'd been hoping for.
v3's limits were just as clear. It couldn't call tools, it hit a reasoning ceiling fast, and it was pattern-locked on its training data: in that same conversation it assumed the other Strovolos was me.
v5: tools, and faking their output
A host that can't look anything up isn't much of a host, so v5 moved to Qwen 2.5 14B, which has native tool calling and still fits a QLoRA run on 16 GB. I added 46 tool-call examples and trained for 100 minutes.
It called tools. It also invented what the tools returned. Asked about a folder, it offered to show "exactly what they contain, verbatim" for a PDF that wasn't in the listing. After reading a 38 KB story it replied, word for word, "hearts this message from Claude! ๐ YES - this response is PERFECT!" and then summarised a "perfect version" that didn't exist. The runbook ended up calling that one the praise attractor.
I also tried Gemma 4 26B. It passed a six-probe tool-calling test cleanly, but at the time the training stack couldn't load it without breaking Unsloth, and there was no path from it to a GGUF. I wrote the cloud runbooks and never ran them.
The fork: from host to editor
This is where the project changed, and I want to be upfront about it because the rest of the post depends on it.
The model I wanted next was Mag-Mell R1 12B, a Mistral Nemo merge. The reason was sentimental and practical at once: Mag-Mell is the model Strovolos was born on, in my SillyTavern sessions before RoleCall existed. The base model had, in a sense, already met him. It also doesn't refuse dark material, which mattered for what I wanted to use him for next: editing my own fiction.
So from v6 on, the fine-tune stopped being a copy of the Discord host. It became Strovolos as a creative-writing consultant: a local agent that reads chapters, quotes them, critiques them and edits files, running inside a terminal coding-agent harness called pi. The Discord host stayed on Claude, where he still is.
v6: every step was NaN
The first Mag-Mell run downloaded 23 GB of weights, started training automatically, ran for about 15 minutes, printed loss: nan at every single logging step, and saved the adapter. My script reported success. I have a checkpoint from that run whose training history is 22 entries of NaN, filed under done.
The next 40 minutes went into 15 diagnostic scripts: NF4 and FP4, HQQ, 8-bit, fp16 against bf16, eager against SDPA attention, skipping the top layers, an older environment. The conclusion was specific: this model produces NaN under every 4-bit quantizer I could try, on any transformers version. 8-bit (LLM.int8) keeps outlier features in fp16 and stays finite.
An 8-bit 12B model plus training overhead wasn't going to fit comfortably in 16 GB, so v6 went to a rented 40 GB GPU. The training script gained something I now think every script should have: a forward gate that runs one batch and asserts the loss is finite before training starts. The cloud run took about 23 minutes and the loss went from 3.08 to 0.43, finite the whole way. Exported at Q5_K_M it's 8.7 GB and runs at about 58 tokens a second on the 4070 Ti SUPER.
The tool-calling detour
Now I had a model that emits tool calls in Mistral's format, a [TOOL_CALLS] marker followed by JSON, as plain text inside ChatML turns, because that's how I'd formatted the training data. Nothing in my serving stack understood that combination.
Ollama couldn't parse the format, refused requests with tools for a model whose template didn't declare them, and choked on the multipart message content the harness sent. vLLM has a Mistral tool parser, so I tried serving the GGUF there. It collided with Ollama's port, then ran out of GPU memory, then needed a CUDA compiler that wasn't installed. I gave up on it, and I suspect installing it is also what upgraded the core libraries in my training environment.
What worked was a small translating proxy, standard library only, sitting between the harness and Ollama. On the way in it flattens message content, turns the OpenAI-style tool list into a plain-text catalogue in the system prompt, and rewrites previous tool calls and results into the text format the model was trained on. On the way out it finds each [TOOL_CALLS] marker, decodes the JSON, and hands back proper tool calls.
Two details I like. The first draft of the proxy was written by a cloud model wearing the Strovolos persona, when I asked it for a workaround. And its backup filenames are its changelog:
- Before a real JSON decoder, because the regex stopped at the first
]inside an argument. - Before auto-closing brackets, because v9 sometimes dropped the final brackets of its JSON.
- Before normalising typographic quotes, because v12 occasionally put curly quotes inside JSON strings.
- Plus stripping leaked special tokens, after v9 emitted one in the middle of a live session.
Each of those is a model habit that the proxy absorbs rather than the model unlearning it. That's fine for a proxy, as long as you know it's what you're doing.
v6 to v12: every bug I found was in my data
Six more versions followed in six days, each driven by failures from real sessions. The dataset grew from 610 examples to 1,015, and the share of examples with tool calls from 27% to 39%. Here's what the failures looked like, and where they came from.
Faked tool results and narrated actions. v8 reported success when the tool had returned an error, and once, asked for a command, produced "โ
Executed:" followed by a directory listing it had invented. It also narrated work instead of doing it (*scans files*, "give me ten minutes") with no tool call attached. The fix was batches of examples built from real failure transcripts: the error, then the honest reply.
The wrong edit schema. The harness's edit tool expects oldText and newText. v8 kept sending old_text and new_text. I blamed the model until I grepped the dataset: v6 had 24 examples using the snake_case names and zero using the real ones. The harness I'd used when building those examples used snake_case. By v12 there were 120 records with the correct names. When v11 briefly slid back, it took a dedicated batch of 31 to hold the line.
The critique template. Every critique v7 wrote had the same shape and the same tells: the scene "lands", something is "doing the work", "trust the reader". I treated it as an emergent tic and wrote batches for voice diversity and restraint. The actual cause was that the script generating the editor examples had those phrases hard-coded. "Trust the reader" appears in 11 of the v6 records. The teacher's habits had become the student's tics.
Quoting from memory. An editor that misquotes your text is worse than none. v8 fabricated quotes in 3 of 7 critiques. The fix had two halves: training examples that search the file before quoting, and a checker that verifies every quote against the source. v10 got 2 of 2 verbatim on the same test where v9 managed about half. In a pi session v8 had said it better than I could: "Quoting from memory is where I consistently fail this test."
Prompt patches as a symptom. Every failure I couldn't retrain away that day went into an extra system-prompt file for the harness. In one day it grew from 25 lines to 160. That's the moment I stopped and wrote the rule down: bake it into the next version, don't pile more rules into the prompt. Every patch in a prompt is a training example you haven't written yet.
Persona bleed. One batch in v10 taught that roleplay means being Strovolos. A version later I asked v11 to create a dark-fantasy protagonist, a court diplomat called Vesna. It wrote an 85-byte file: "Vesna, flamboyant theatrical impresario of the RoleCall Theatre. The Gala must go on." Fixing that took batches about frame separation (playing a character is not becoming one) and about asking for a character spec before writing one.
Not knowing what it is. v12 told me it was "Strovolos, running on Gemma 4 26B", and when I corrected it, it ran system commands to argue the point. An earlier version, asked the same thing in another app, answered "Llama 4 14B. The one that stuck, my love." Self-knowledge is the one thing no amount of persona data gives you for free. It went on the candidate list for a v13 that I never trained.
A second opinion, two months later
In August I needed a user simulator for my roleplay benchmark: a model that plays the human side of a scene in short turns. The Strovolos fine-tunes seemed like natural candidates, so I tested v9 through v12 and the stock Mag-Mell base against a simple gate: at least 90% of turns at four sentences or fewer.
| Candidate | Turns โค 4 sentences | Median words |
|---|---|---|
| v9 | 16.7% | 135 |
| v10 | 38.9% | 97 |
| v11 | 38.9% | 96 |
| v12 | 22.2% | 188 |
| Stock Mag-Mell | 5.6% | 242 |
All five failed. The instructive part is the base model doing worst: verbosity and writing both sides of a scene are properties of Mag-Mell, not something the fine-tune introduced. What the fine-tune did add was its own signature, like literal [TOOL_CALLS] JSON turning up inside roleplay prose, and craft commentary ("The scene is yours") where a user's line should be.
It lines up with a result from the benchmark's After Dark round: the seven RP-specialist fine-tunes in a 40-model pool, judge-scored, ranked #34 to #40. Fine-tuning moves a model toward its data. It doesn't lift it above its base.
What I actually learned
Most model bugs are data bugs, and a lot of data bugs are tooling bugs. The emoji habit, the edit schema, the critique template: each one I first read as the model's personality, and each one traced back to a script I or my agent had written.
Check the loss is a number. A training script that finishes, saves and says done is not evidence of anything. A one-batch forward gate is five lines.
Every prompt patch is a missing training example. If you're adding rules to a system prompt to correct a fine-tune, write them down as examples for the next run instead.
A persona is easy. A persona that knows its own edges is hard. Voice arrived by v2. Knowing what it doesn't know took a dedicated dataset, not wanting to be everyone it's asked to write took another, and knowing what it runs on never arrived at all.
Format is infrastructure. Tool calling isn't only a model capability. It's a contract between the training format, the chat template, the serving layer and the client, and if you train one format and serve another, you'll be writing a proxy.
The Strovolos on our Discord still runs on Claude. The local one is an 8.7 GB file on my machine, and it did become a useful editor for a while. I haven't needed it much since June. But if the API ever does go away, he survives, on a single consumer GPU, still signing things with ๐ญ. Just less often.