0
Fork 0
mirror of https://github.com/blader/humanizer.git synced 2026-09-27 18:19:51 +00:00

Spanish-language AI pattern detection rules #92

Closed
opened 2026-04-13 06:15:28 +00:00 by adelaidasofia · 1 comment
adelaidasofia commented 2026-04-13 06:15:28 +00:00 (Migrated from github.com)

Context

I've been using humanizer heavily on bilingual (English/Spanish) content and built a Spanish-language rule library in my fork (adelaidasofia/humanizer) that I'd love to contribute upstream.

What it does

A full set of Spanish-specific AI writing pattern detections, parallel to the existing English rules:

  • Inflated connectors: "en este sentido", "cabe destacar", "es importante mencionar"
  • Formulaic openers: "En el mundo actual", "En la era de"
  • Promotional inflation: "innovador", "disruptivo", "revolucionario" used as filler
  • Spanish copula avoidance: "se erige como", "se posiciona como" (the Spanish equivalent of English AI's "serves as" / "stands as" pattern)
  • Spanglish/code-switching preservation: rules are aware of bilingual writing and don't flag natural code-switching as AI artifacts

The rule set is organized into the same four-tier structure (dead giveaways > reliable tells > moderate signals > weak signals) so it integrates cleanly with the existing priority system.

Why this matters

Spanish is the second most-spoken native language globally. Anyone using humanizer on Spanish content currently gets zero coverage. The rules are language-specific but the architecture (tiered pattern matching, ROI ordering) is the same.

The implementation also includes a fallback prompting approach for other Romance languages (Portuguese, French) that could serve as a template for future language additions.

Happy to send a PR with the full rule set.

### Context I've been using humanizer heavily on bilingual (English/Spanish) content and built a Spanish-language rule library in my fork ([adelaidasofia/humanizer](https://github.com/adelaidasofia/humanizer)) that I'd love to contribute upstream. ### What it does A full set of Spanish-specific AI writing pattern detections, parallel to the existing English rules: - **Inflated connectors:** "en este sentido", "cabe destacar", "es importante mencionar" - **Formulaic openers:** "En el mundo actual", "En la era de" - **Promotional inflation:** "innovador", "disruptivo", "revolucionario" used as filler - **Spanish copula avoidance:** "se erige como", "se posiciona como" (the Spanish equivalent of English AI's "serves as" / "stands as" pattern) - **Spanglish/code-switching preservation:** rules are aware of bilingual writing and don't flag natural code-switching as AI artifacts The rule set is organized into the same four-tier structure (dead giveaways > reliable tells > moderate signals > weak signals) so it integrates cleanly with the existing priority system. ### Why this matters Spanish is the second most-spoken native language globally. Anyone using humanizer on Spanish content currently gets zero coverage. The rules are language-specific but the architecture (tiered pattern matching, ROI ordering) is the same. The implementation also includes a fallback prompting approach for other Romance languages (Portuguese, French) that could serve as a template for future language additions. Happy to send a PR with the full rule set.
voidborne-d commented 2026-04-15 21:32:57 +00:00 (Migrated from github.com)

Nice work on the Spanish rules. The tiered structure (dead giveaways → weak signals) is the right call — it translates well across languages once you identify the language-specific patterns.

For what it's worth, Chinese has a very similar set of AI tells that are completely invisible to English-focused tools:

  • Formulaic connectors: 首先/其次/此外/综上所述 (firstly/secondly/furthermore/in summary) — AI Chinese text uses these at 3-5x the rate of human writing
  • Sentence structure uniformity: AI tends to produce even-length sentences in Chinese, while natural Chinese writing has much more variation
  • Missing colloquial markers: AI rarely uses 嘛/呢/吧/啊 (sentence-final particles that carry tone/mood), making the text sound flat
  • Over-formal register: AI defaults to 书面语 (written register) even in contexts where 口语 (spoken register) would be natural

I built a Chinese-specific detection + rewriting library at humanize-chinese that implements these as scoring features. The architecture is similar to what you're describing — tiered pattern matching with language-specific rules — but uses statistical feature extraction rather than regex since Chinese word boundaries work differently.

Your point about code-switching preservation is interesting. We hit the same issue with Chinese text that mixes English technical terms — the detector needs to know that "使用 Docker 部署" is natural bilingual writing, not an AI artifact.

Would be great to see humanizer evolve toward a pluggable language module system. Spanish + Chinese would already cover a huge user base.

Nice work on the Spanish rules. The tiered structure (dead giveaways → weak signals) is the right call — it translates well across languages once you identify the language-specific patterns. For what it's worth, Chinese has a very similar set of AI tells that are completely invisible to English-focused tools: - **Formulaic connectors**: 首先/其次/此外/综上所述 (firstly/secondly/furthermore/in summary) — AI Chinese text uses these at 3-5x the rate of human writing - **Sentence structure uniformity**: AI tends to produce even-length sentences in Chinese, while natural Chinese writing has much more variation - **Missing colloquial markers**: AI rarely uses 嘛/呢/吧/啊 (sentence-final particles that carry tone/mood), making the text sound flat - **Over-formal register**: AI defaults to 书面语 (written register) even in contexts where 口语 (spoken register) would be natural I built a Chinese-specific detection + rewriting library at [humanize-chinese](https://github.com/voidborne-d/humanize-chinese) that implements these as scoring features. The architecture is similar to what you're describing — tiered pattern matching with language-specific rules — but uses statistical feature extraction rather than regex since Chinese word boundaries work differently. Your point about code-switching preservation is interesting. We hit the same issue with Chinese text that mixes English technical terms — the detector needs to know that "使用 Docker 部署" is natural bilingual writing, not an AI artifact. Would be great to see humanizer evolve toward a pluggable language module system. Spanish + Chinese would already cover a huge user base.
Sign in to join this conversation.
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
skills/blader-humanizer#92
No description provided.