0
Fork 0
mirror of https://github.com/blader/humanizer.git synced 2026-09-27 18:19:51 +00:00

Feature Request: Sentence Starter Variation (Pronoun Repetition) #206

Closed
opened 2026-08-02 16:36:07 +00:00 by mrman1717 · 1 comment
mrman1717 commented 2026-08-02 16:36:07 +00:00 (Migrated from github.com)

Description of the Problem

When generating or processing narrative text, AI models frequently fall into a repetitive syntactic loop where consecutive sentences begin with the exact same word. While this pattern can occur with any overused sentence starter, it appears to mostly apply to the pronouns "she" and "he". This repetitive structure creates a robotic, stilted rhythm that immediately breaks narrative immersion and flags the text as artificially generated.

Examples

In a recent analysis of a generated narrative text file from a 2-hour-long audiobook, the pronoun "she" was heavily overused as a sentence starter:

Total occurrences as a sentence starter: 147
Occurrences happening twice in a row: 34
Occurrences happening three times in a row: 11

A specific example of this repetitive stacking from the text:

"She noted the door. She noted the lock on it. She filed both away."

Potential Solutions

To humanise the text and break up this unnatural pacing, the following interventions could be implemented into the skill:

  • Sentence Consolidation: Merge short, choppy sentences using conjunctions to improve flow (e.g., transforming the example above into "She noted the door and its lock, filing both away for later.").
  • Subject Shifting: Rework the syntax to make the object or environment the subject rather than the character (e.g., "The door and its lock caught her attention.").
  • Action-First Restructuring: Introduce gerunds, participle phrases, or adverbial clauses at the beginning of the sentence to delay the pronoun (e.g., "Noting the lock on the door, she filed the information away.").
  • Algorithmic Penalty: Introduce a threshold that flags the generation process if the exact same starting token is used more than twice in succession, prompting an automatic syntactic rewrite for variety.

Implementing this pattern would significantly improve the natural cadence and readability of generated prose.

### Description of the Problem When generating or processing narrative text, AI models frequently fall into a repetitive syntactic loop where consecutive sentences begin with the exact same word. While this pattern can occur with any overused sentence starter, it appears to mostly apply to the pronouns "she" and "he". This repetitive structure creates a robotic, stilted rhythm that immediately breaks narrative immersion and flags the text as artificially generated. ### Examples In a recent analysis of a generated narrative text file from a 2-hour-long audiobook, the pronoun "she" was heavily overused as a sentence starter: Total occurrences as a sentence starter: 147 Occurrences happening twice in a row: 34 Occurrences happening three times in a row: 11 A specific example of this repetitive stacking from the text: "She noted the door. She noted the lock on it. She filed both away." ### Potential Solutions To humanise the text and break up this unnatural pacing, the following interventions could be implemented into the skill: - **Sentence Consolidation:** Merge short, choppy sentences using conjunctions to improve flow (e.g., transforming the example above into "She noted the door and its lock, filing both away for later."). - **Subject Shifting:** Rework the syntax to make the object or environment the subject rather than the character (e.g., "The door and its lock caught her attention."). - **Action-First Restructuring:** Introduce gerunds, participle phrases, or adverbial clauses at the beginning of the sentence to delay the pronoun (e.g., "Noting the lock on the door, she filed the information away."). - **Algorithmic Penalty:** Introduce a threshold that flags the generation process if the exact same starting token is used more than twice in succession, prompting an automatic syntactic rewrite for variety. Implementing this pattern would significantly improve the natural cadence and readability of generated prose.
somtri commented 2026-08-03 06:36:50 +00:00 (Migrated from github.com)

I'd like to take this one.

Rather than adding a numbered pattern, I think this belongs as an extension to an existing section. You closed #171 with the note that a new numbered pattern carries permanent prompt and maintenance cost, and that the narrow idea might still fit as a small extension to an existing rule. Repeated sentence-initial subjects look like the same case to me, so the diff would be a few lines on the closest existing section plus one false-positive guard in Detection Guidance, leaving the count at 33.

I'd skip the fourth remedy in the description. A threshold that counts starting tokens needs code to run in, and this repo ships a Markdown prompt.

The guard is the part I want to get right. Deliberate anaphora is an old device and plenty of good prose repeats a sentence opening on purpose, so the tell can't be "same first word twice." It has to be a run of them where the repetition isn't doing any work.

I'd like to take this one. Rather than adding a numbered pattern, I think this belongs as an extension to an existing section. You closed #171 with the note that a new numbered pattern carries permanent prompt and maintenance cost, and that the narrow idea might still fit as a small extension to an existing rule. Repeated sentence-initial subjects look like the same case to me, so the diff would be a few lines on the closest existing section plus one false-positive guard in Detection Guidance, leaving the count at 33. I'd skip the fourth remedy in the description. A threshold that counts starting tokens needs code to run in, and this repo ships a Markdown prompt. The guard is the part I want to get right. Deliberate anaphora is an old device and plenty of good prose repeats a sentence opening on purpose, so the tell can't be "same first word twice." It has to be a run of them where the repetition isn't doing any work.
Sign in to join this conversation.
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
skills/blader-humanizer#206
No description provided.