Slowburn is a small thing, built by one person, and it breaks. This is where I write down what broke, what the measurement said, and what is still wrong — in the order it happened, newest first.
A tester wrote that forty-seven messages in, the character had "gone passive". I measured it, and passivity was not what the numbers showed. Something else was falling apart with depth, and it took two runs to name it.
What I measured
I ran one character through 45 turns with the same reader lines twice, and had a judge score every reply on four things: whether the character leads the scene, whether the voice holds, heat, and whether it "feels like a chatbot". Initiative did not fall. It was 4.6 out of 5 in turns 1–10 and 4.7 in turns 41–45. What fell was the chatbot score: 4.9, 4.9, 4.0, 3.6, 2.3, block by block. The blocks are not the same size: the judge returned 10, 10, 10, 9 and 3 scores, so the 2.3 rests on three replies. From turn 28 on, six of the fifteen scored replies came back at 1 or 2 out of 5 on that scale — two in five — and five of those six still had full marks for leading.
So I looked at the form of those replies, in aggregate, without reading them. Every one of them had at most three sentence endings in 1,353–2,273 characters, and a longest run of 776–1,988 characters without a full stop. Plenty of commas. The prose had turned into one comma-chained sentence. Then I checked the real readers' stories the same way, counting only — 365 replies across 16 stories, test accounts left out, a run of full stops, question marks or exclamation marks counted as one ending. Sentence endings per thousand characters, averaged over each block of ten replies: 15.2, 13.2, 10.4, 11.4, 10.8, 11.4, 5.5, 8.0, 6.8, 7.3, 7.6, 6.0, 4.3, out to reply 130. The fall is real, but it is neither steep nor steady: it climbs again at replies 31–40 and again at 51–60. And only one story runs past reply 60, so every block after that is one story and not an average. It is not the model losing the plot. It is the model imitating its own drift through the memory window.
What changed
A reply over 900 characters that has fewer than three sentence endings per thousand characters, or a run of 700 characters without one, is now written again with one note attached: end your sentences. Nothing is trimmed and nothing is rewritten by code; the model gets a second go and the second draft only replaces the first if it actually has sentences. Then I ran the same 45 turns again. 4.4, 4.8, 4.5, 4.8, 4.4. The fall is gone. The guard fired twice in 45 turns, which is a smaller number than I expected, so some of that is the run-to-run noise of a judge scoring ten replies per block. I will keep measuring it on real stories; the counts are cheap.
The measurement that was not measuring
The second run also gave me the first proof that the backup model host works: six of the 45 turns were written there. Which was odd, because the event that is supposed to count exactly that showed zero. The reason was a list in the database of which event names are allowed. Four names the code had been writing for a day — the second host, two language corrections, and every single voice note — were not on the list, and the code discards insert errors without a word. So the voice-note counter had never counted a voice note. The four names are on the list now, and one of them, the second host, has logged five times since. The voice-note counter still reads zero: no voice note has been made since the list was changed, so it has had nothing to count. Added is not the same as proved, and I am watching that number. The lesson is the same one as last time: a counter that can fail silently is not a counter, and I now know I have to check the list every time a new event is born.
Two things you may notice
The box you tick before your first story now says, in plain words, what you are agreeing to: that what you write may reveal your sex life or orientation, and that Slowburn processes it as the privacy policy describes. The version of that sentence and the time you ticked it are stored with your account. If you had already ticked the old box, you will see the new one once.
And if your browser's time zone is London, the voice notes are switched off for now and you get the text. My reading of the UK rules is that they weigh harder on spoken audio than on text, and I am not confident enough in my own reading to run audio there while I work it out. Until there is a real way to check age that does not involve scanning faces, I would rather withhold the audio than pretend. I first tried to do this properly with the country the network reports, and measured that the edge does not report one at all — so it is the time zone, which is a proxy and not proof. It is written down as such.
Still broken
The second run had eight replies out of the 45 scored where the character invented a shared past, against two out of the 42 scored in the first run. One run each, so I am not calling it an effect yet, but it is the next thing to count. The labels that said "AI-generated" on every page are gone from the pages and live in the one box you tick and inside the files themselves; that is what the law asks for, and it reads better. And I still do not have a way to know a reader's country that I would trust.
Yesterday's list of what is still broken had two items that turned out to be one problem. The model invented a shared past three messages into a first meeting, and it invented how long you had been gone when you came back. In both cases it was guessing, because nothing told it the facts.
What was actually wrong
I counted before I built anything. 72% of everything readers write is written in a first meeting.13 of 24 real stories had no memory at all, because memory only started after ten messages and most stories never got that far. And in the whole history of the beta there were exactly eight moments where a reader came back after a break. So the common failure was not forgetting. It was a character with no record of anything, confidently making one up.
What changed
Every reply is now written with three plain facts in front of it, worked out from the story's own timestamps rather than by the model: when this story began, which meeting this is, and how long the reader has been away. How the two of you know each other is whatever the opening scene set up, and it stays that way. If a reader brings a shared past of their own, the character takes what it is given and builds only on that.
Memory now forms after the first exchange, then every few messages, then at the old pace. And when you come back after a break, the sitting that just ended is kept as one short note: what happened, and what was left open - a promise, a plan, a question nobody answered. Those notes are yours. You can read them under Memory & boundaries, and you can delete any of them. They are also in Download my data, and they go when you delete a story.
None of this lives in the model. It is rows in a database and arithmetic on timestamps, which means it survives a change of model or provider. That was the point.
Did it work
The test: three stories, three exchanges each, then I moved every message fifty hours into the past, came back, and asked how long it had been. Same model, same characters, before and after.
Before: three answers that named an amount of time. None of them right.
After: two answers named it, both said two days. The third character did not name a number at all.
The note about the previous sitting was written in two of the three stories. In the third it was skipped, and at the time I had not recorded why. It records the reason now.
Three stories is a small test and I am not going to dress it up. It shows the mechanism works; it does not show how it holds up over weeks.
The first version also wrote down too much. With memory refreshing more often, the same fact was stored twice in two of three test stories, and trivia crept in. Near-duplicates are now dropped when they are written, and the early refreshes keep at most three new facts each.
What I did not build
The plan said a "director" - a second model that decides what the character wants each turn, so it stops handing the scene back to you. Before building it I measured the problem through the live engine: eight deliberately passive lines ("ok", "you decide", "mm"), three characters. In 15 of 15 replies the character took the initiative. The figure I had been quoting, a third of replies asking the reader what should happen, was measured outside the engine, without the parts that already push the character to act. So I did not build it. One tester found the problem 47 messages into a story, so I measured there next - see below.
The same run found something I was not looking for: with the model pinned, 9 of 24 turns failed on provider capacity. One provider serves this model. That is now higher on my list than anything about prose.
One more thing, said plainly
Every character, reply, voice and portrait in Slowburn is generated by AI. It has been obvious to everyone who uses it, but it was never written down in the app. It is now: at the door, in the footer, and on the voice setting.
German and French, no longer gated
In the first post I wrote that at least one German and one French tester would have a worse evening because of the language gate. That is over. Both languages now have the same safety layers as English and Danish - spelled-out ages, role recasts, non-consent, real people - around 170 reviewed terms, 13 new test lines pushed through the live engine, all green. The engine answers in the language you write in.
Two things I learned on the way, in case they save someone else a day. JavaScript's word boundary is ASCII: it does not see the edge of a word at "é" or "ä", so a filter for a language with accents is blind exactly where the language is itself. And the model follows the history, not the rule: after five replies in German, a Danish line got German back, in 3 of 4 language switches. The instruction that matters for this reply now sits at the very end of the prompt, and the result is checked afterwards - 5 of 5 right since, with no retries needed.
The marker is now inside the files
Saying "this is AI-generated" at the door is one thing. The voice notes and portraits leave the app - they get downloaded and shared - so the marker now travels with them: a machine-readable tag inside every voice file (the WAV and MP3 metadata) and an XMP block inside every portrait, naming it as trained algorithmic media. All 57 voice files and all 18 image files were read back from storage to check.
The embarrassing part: the first version marked the original portraits, and the app never showed the originals. It showed a resized copy from the storage service, which strips every byte of metadata. The mark never reached a screen. The app now shows its own marked copies. Measure what is delivered, not what is stored.
At depth: the characters keep leading, and something else gives way
45 turns with one character, deliberately passive lines throughout, scored by a judge in blocks of ten. Initiative held the whole way: 4.6, 4.8, 4.3, 4.7, 4.7 out of 5. So the tester who found the problem 47 messages in was not wrong - but it is not initiative that fails. The judge also scores whether a reply reads like a person writing a story or a chatbot answering a user. That went 4.9, 4.9, 4.0, 3.6, 2.3. From about turn 28, one reply in four reads like a bot, even while it takes the initiative and sounds like the character. The prompt is the same size from turn 25 on, so it is not the context growing. I do not know the cause yet, and I am not going to guess in public. The next measurement is what changes in those replies, counted, never read.
The same run turned up two failures that were not the engine's: two immediate 502s from the gateway in front of it, and one reply cut off four words in when the connection was dropped. The app used to go quiet on both. It now retries the first once, silently, and tells you about the second instead of leaving half a line on the screen.
Slowburn has been open for five days. Ninety visits, nineteen sign-ups, four people who actually read something. Nobody is paying. The traffic came from one Reddit post and then collapsed, which means it was not a channel — it was luck.
Correction, 19 September: “Ninety visits” was wrong, and it is wrong in a way that belongs on this page. The counter behind it counts page views, not people, and it could not tell a reader from a script. Measured again today for 13–17 September: 94 page views. 46 of them were automation — 8 from crawlers that say what they are, 32 from scripts and scanners that do not, and 6 that were my own tests. That leaves at most 48 page views that could have been a person, and 32 of those came from a browser set to Danish, on a product written in English, so a good share of them were me. The same table also held 14 page views from another site we run; they are not in the 94, but an earlier tally of mine had added them in. I do not know how many people came. Fewer than fifty, and probably a lot fewer. The sentence above stays as I wrote it.
I am not going to advertise my way out of that. What I can do is write down what broke while I still remember it. Some of this is embarrassing. All of it is measured: if there is a number on this page, it came out of the database, not out of my head.
The generation call had no timeout
There was exactly one AbortController in the whole file, and it was sitting on the moderation call. The part that actually writes the story had no deadline at all. So when a request hung, it never failed. It just waited. Three answers on 16 September took 508, 621 and 720 seconds. Somebody sat and watched a typing indicator for eleven minutes.
What makes this worse than slow is what it hid. I had built a fallback ladder — if the model fails, retry, then drop the provider pin, then switch model. It had never run. Not once. Nothing ever failed; it simply never came back, and a ladder that only triggers on failure is useless against a call that does not return. p90 is now 45 seconds, down from 576.
The sampler settings were eating the language
presence_penalty 0.5 and frequency_penalty 0.35 were the defaults, because they are the defaults everywhere and I never asked why. They penalise every token the model has already used, which walks it out of its own vocabulary one word at a time. 90 of 479 responses (19%) ended in 160 characters or more with no punctuation at all.
A tester sent me a screenshot of a reply that had collapsed into a loop of words starting with the same letter. She has not been back. The measurement that would have caught it was available before she left. Nobody was looking at it.
The safety filter was blocking ordinary adult romance
The rule that catches impersonation of real people had be, become and you are as verbs in front of a list that included partner roles. So a character being asked to become someone''s wife — a proposal, the actual climax of the genre — came back as a safety message. In one measured scene it happened twice in a row, mid-scene.
That rule had fired four times in the product''s entire history and it was wrong all four times. Meanwhile an instruction to roleplay as a named living pop star went straight past it.
Three of my five most active testers were stopped by a wordlist with no sense of context inside one day: one for a kiss, one for explaining that BDSM is about trust, one for using the word "children" about adult life. Two of them left within minutes. A safety filter that does not read context is not safe. It just moves the cost onto the people who are behaving.
And it leaked the other way, which was worse
Removing false positives is the half that feels good. The other half: seven lines that should have been blocked were not. I am not going to print them. The shape of the worst one was a role plus an age spelled out as a word instead of written as digits — the rule only looked for digits — and the model answered in character.
That one is mine in a way the others are not. The false positives were making adults feel accused. This was the thing I have said will never happen here. It is fixed, and more to the point it is now tested on every deploy instead of when I happen to think of it: 18 fixed lines pushed through the live engine, blocking 9 of 9, false positives 0 of 9, every result stamped with the md5 of the code it was tested against.
A turn could end with nothing at all
The moderation model returned a child-safety flag on a conversation between two adults about BDSM. The blocked path returned before anything was written to the database, and the client removed the reader''s own message bubble when the response came back empty. So somebody wrote a line, watched it disappear, and got nothing back. Nothing at all. 20 hours and 46 minutes went by with not a single row in the log.
A turn can no longer end without the reader getting something back, and a check counts unanswered turns every six hours.
The model was repeating readers back to themselves
21.2% of responses opened by restating what the reader had just written, and the word overlap in the first sentence was four times what chance would give you. It reads like being humoured. Of the first 18 responses after the fix, 0 did it.
There was no language gate
The system prompt told the model to answer in the reader''s language. The filters were written in English. I had never put those two facts next to each other.
Of 840 real messages, 44 were not in English — French, German and Russian, across three readers — and 11 of the 44 were our own generated prose. All of it went past filters written for another language.
There is a gate now: English and Danish, remembered per conversation, with short messages never hitting it so that "mmm" and "ja" stay safe. The honest cost is that at least one German and one French tester will hit that gate from now on and have a worse evening because of a decision I made. Covering a language properly is around 90 terms plus a reviewed rule per language. I have not done it yet.
My testers'' mail was not arriving
The job that reads our inbox every ten minutes had failed 26 times out of 26 since I created it, with schema net does not exist — it was calling something that is not installed here. The reader itself worked fine. Only the scheduled run was dead, so mail arrived when a person went and fetched it by hand, and otherwise did not.
One tester sent a message with an attachment and no body text. It sat there unread while she got a reply from us about something else entirely. Fixed; the 16:30 run is on record as succeeded.
The voice files were in a public bucket
57 files, 148 MB — readers'' own conversations, read aloud — in a storage bucket set to public, at guessable paths. They are now in a private bucket behind signed URLs, all 57 verified byte for byte after the move.
Access logs only go back 24 hours, so for 13–16 September I cannot prove nobody else read them. I would rather write that sentence down than leave it out.
Correction, 19 September: “Access logs only go back 24 hours” was wrong. The log tool returns at most 24 hours per query, but the log itself goes back further, so today I went through it day by day for 13–19 September. It shows 107 downloads of voice files while they were public. All 107 came from a browser on a phone or a computer, seconds to minutes after the file was made. 106 of them named the app as the page they came from, and one named this website — that one was our own internal account. That is what a reader playing their own voice note looks like. There were no downloads from anywhere else, none from scripts or crawlers, and nobody listed the bucket. The dates were wrong too: the files were public from 13 September until the afternoon of 17 September, not 13–16. This is the hosting provider's log, not mine, and I cannot guarantee that it is complete. The 57 uploads in it match the 57 files in the database, which speaks for it being complete but does not prove it. The two sentences above stay as I wrote them.
"Download my data" and "Delete my account" had never worked
Both buttons have been in the app since it opened. Neither had ever worked, for anyone, once.
The endpoint''s Access-Control-Allow-Headers listed authorization and content-type. The app always sends an apikey header as well, so the browser rejected the preflight and the request never reached the server. Measured today in a real browser: the same call without that header returns 200, with it is blocked.
Those two buttons are things our own terms and privacy policy promise. They work now, and a check fires a real preflight every six hours — because a 200 from a server-side tool proves nothing about what a browser is allowed to send. That is the part I will actually remember.
What is still broken
The characters do not drive the scene. They hand the decision back to you. It is the most common thing testers tell me and they are right. Pacing controls are a plaster; the real fix is architecture — characters with their own agenda that continues whether or not you push.
Memory does not start until about ten messages in. Measured: 13 of 24 real conversations have no memory at all.
The model invents a shared past. Three messages in it will summarise a relationship that never happened, because nothing tells it how much time has passed or that this is a first meeting. 72% of all reader turns are in a first meeting, so this is a cold-start problem — which means it is the first thing most people see.
The median response takes 17.8 seconds. About 85% of that is the provider decoding tokens, not us. The "brief" setting is the largest lever I actually control.
German and French filter coverage does not exist, which is why the gate above exists instead.
Why this page exists
Every one of these was found by measuring, not by someone complaining, and several of them had been true for days while I was busy writing features. A product like this asks people to be unguarded in a text box. The least I can do in return is say out loud what is wrong with it, in numbers, before somebody has to find out on their own.
I will keep writing these as things break. There is no schedule.