An AI roleplay platform for creative storytellers, built with Chi and Nemo and open to everyone since June 2026, with a free tier and paid plans. Every reply is an agentic turn: a story director planning behind your back, a narrator with tools, lore that files itself, characters with their own minds, and a world you can rewind.
Dual-stream AI architecture — the story director plans, the narrator writes, and neither sees the other's full context
Shadow world simulation — invisible arcs, hidden events and off-screen NPC actions the player discovers organically
Immersion engine — combat, relationships, inventory, quests and 20+ more trackers driven by AI directives
Visual novel mode — authored story graphs, sprites and endings, performed live by the model on a stage
Character Minds — per-character memory, mood, drives and theory of mind on a pgvector semantic layer
Turn-wide undo — deleting a message rewinds every write that turn made, across lore, plans, trackers and minds
Eleven fine-tuned versions of the RoleCall Discord host in eight days, on a single 16 GB GPU plus rented cloud time: from a Llama 3.1 8B persona clone to a Mag-Mell 12B writing editor with tool calling, served through a translating proxy.
Datasets grew from 28 to 1,015 examples, with tool calls rising to 39% of records
Diagnosed a model that produces NaN under every 4-bit quantizer; moved to 8-bit LoRA with a finite-loss forward gate
Traced 'model tics' back to the data generators: a wrong edit schema and hard-coded critique phrases
Built a tool-call proxy between OpenAI-style clients and Mistral-format tool calls in ChatML
An intelligent Retrieval-Augmented Generation system for SillyTavern that gives AI characters long-term memory through semantic search. Instead of relying on fixed context windows, VectHare lets characters recall relevant information from hundreds of messages back — featuring natural memory decay, conditional activation, and multiple vector backends.
Temporal decay — memories naturally fade with configurable half-life, important scenes can be marked as immune
Conditional activation — chunks trigger based on 28 emotion types, keywords, recency, with AND/OR logic
Hybrid search — combines semantic similarity with keyword matching for more accurate retrieval
Multiple backends — built-in Vectra for zero-dep setup, LanceDB for millions of vectors, Qdrant for enterprise
A specialized evaluation framework that measures roleplay quality in large language models across 27 dimensions. It tests what actually matters in interactive roleplay (character consistency, agency respect, lorebook integration, prose craft and genre skills) and calibrates its LLM judges against blind human votes.
27 evaluation dimensions across 3 tiers — from agency respect and continuity to atmospheric dread and structural comedy
Adversarial multi-turn sessions — 12 turns against a user simulator, with seeds that bait specific failure modes
Public blind-vote arena — thousands of community votes, later moved into PlotLight as PlotPoints
Charm vs stamina — single-message and multi-turn human rankings invert (Spearman ρ = −0.24)
Open dataset at lazyweasel/roleplay-bench on HuggingFace, 2,800+ downloads