mirror of
https://github.com/blader/humanizer.git
synced 2026-09-27 18:19:51 +00:00
Feature Request: Sentence Starter Variation (Pronoun Repetition) #206
Labels
No labels
bug
documentation
duplicate
enhancement
good first issue
help wanted
invalid
question
wontfix
No milestone
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
skills/blader-humanizer#206
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Description of the Problem
When generating or processing narrative text, AI models frequently fall into a repetitive syntactic loop where consecutive sentences begin with the exact same word. While this pattern can occur with any overused sentence starter, it appears to mostly apply to the pronouns "she" and "he". This repetitive structure creates a robotic, stilted rhythm that immediately breaks narrative immersion and flags the text as artificially generated.
Examples
In a recent analysis of a generated narrative text file from a 2-hour-long audiobook, the pronoun "she" was heavily overused as a sentence starter:
A specific example of this repetitive stacking from the text:
Potential Solutions
To humanise the text and break up this unnatural pacing, the following interventions could be implemented into the skill:
Implementing this pattern would significantly improve the natural cadence and readability of generated prose.
I'd like to take this one.
Rather than adding a numbered pattern, I think this belongs as an extension to an existing section. You closed #171 with the note that a new numbered pattern carries permanent prompt and maintenance cost, and that the narrow idea might still fit as a small extension to an existing rule. Repeated sentence-initial subjects look like the same case to me, so the diff would be a few lines on the closest existing section plus one false-positive guard in Detection Guidance, leaving the count at 33.
I'd skip the fourth remedy in the description. A threshold that counts starting tokens needs code to run in, and this repo ships a Markdown prompt.
The guard is the part I want to get right. Deliberate anaphora is an old device and plenty of good prose repeats a sentence opening on purpose, so the tell can't be "same first word twice." It has to be a run of them where the repetition isn't doing any work.