mirror of
https://github.com/blader/humanizer.git
synced 2026-09-27 18:19:51 +00:00
Spanish-language AI pattern detection rules #92
Labels
No labels
bug
documentation
duplicate
enhancement
good first issue
help wanted
invalid
question
wontfix
No milestone
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
skills/blader-humanizer#92
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Context
I've been using humanizer heavily on bilingual (English/Spanish) content and built a Spanish-language rule library in my fork (adelaidasofia/humanizer) that I'd love to contribute upstream.
What it does
A full set of Spanish-specific AI writing pattern detections, parallel to the existing English rules:
The rule set is organized into the same four-tier structure (dead giveaways > reliable tells > moderate signals > weak signals) so it integrates cleanly with the existing priority system.
Why this matters
Spanish is the second most-spoken native language globally. Anyone using humanizer on Spanish content currently gets zero coverage. The rules are language-specific but the architecture (tiered pattern matching, ROI ordering) is the same.
The implementation also includes a fallback prompting approach for other Romance languages (Portuguese, French) that could serve as a template for future language additions.
Happy to send a PR with the full rule set.
Nice work on the Spanish rules. The tiered structure (dead giveaways → weak signals) is the right call — it translates well across languages once you identify the language-specific patterns.
For what it's worth, Chinese has a very similar set of AI tells that are completely invisible to English-focused tools:
I built a Chinese-specific detection + rewriting library at humanize-chinese that implements these as scoring features. The architecture is similar to what you're describing — tiered pattern matching with language-specific rules — but uses statistical feature extraction rather than regex since Chinese word boundaries work differently.
Your point about code-switching preservation is interesting. We hit the same issue with Chinese text that mixes English technical terms — the detector needs to know that "使用 Docker 部署" is natural bilingual writing, not an AI artifact.
Would be great to see humanizer evolve toward a pluggable language module system. Spanish + Chinese would already cover a huge user base.