0
Fork 0
mirror of https://github.com/blader/humanizer.git synced 2026-09-27 18:19:51 +00:00

i18n: Simplified Chinese (zh-CN) adaptation as a separate repo #203

Closed
opened 2026-08-01 14:46:48 +00:00 by jiji262 · 0 comments
jiji262 commented 2026-08-01 14:46:48 +00:00 (Migrated from github.com)

Hi @blader — following #163 (French), #138 (Spanish) and #194 (Traditional Chinese), I built a Simplified Chinese (zh-CN) localization of humanizer. Flagging it here for discoverability, not to request a merge.

Repo: https://github.com/jiji262/humanizer-chinese

You noted in #163 that you'd rather localized variants stay as separate community repositories so each language can evolve without adding duplicate runtime authorities to the core project. This one fits that shape: standalone repo, MIT, with blader/humanizer attribution kept in the LICENSE, README, and SKILL.md.

Why it is an adaptation, not a translation

Three of your rules invert or soften in Simplified Chinese, so a literal port would actively damage correct prose:

  • Curly quotes (§19) invert. “” are the standard in Simplified Chinese (GB/T 15834-2011), not a tell. The rule is replaced with detection of full-width/half-width punctuation mixing, which is the real machine fingerprint.
  • The em dash ban (§14) is softened, not copied. The Chinese dash —— is legitimate punctuation for explanation and topic shifts. Only overuse, paired parenthetical dashes, and mis-typesetting (single-width —, or - / -- standing in for a dash) are flagged.
  • Subjectless sentences (§13) are idiomatic Chinese, not hidden-actor passives. Flagging them would gut normal writing, so that slot instead targets Europeanized 被-constructions and translationese.
  • Hyphenated word pairs (§26) are dropped — Chinese has no hyphenated compounds.

Net: 24 patterns map one-to-one, 8 collapse into 4, 1 is dropped, and 8 Chinese-specific patterns are added, for 36 total.

Chinese-specific patterns, with sources

  • Lecture-note signposting (首先…其次…最后, 值得注意的是, 综上所述) — AI conjunction density measured at roughly 3× human in a 6,586-text parallel corpus study (CCL 2023, Beijing Language and Culture University).
  • Parallelism and antithesis stacking — measured at roughly 6× human density in Chinese AI text, the strongest single quantitative tell found.
  • Bureaucratic empty verbs (对…进行分析, 作出调整) — the Europeanized-Chinese problem Yu Kwang-chung criticized in 1987, which LLM training data amplifies.
  • Uniformly written register — AI modal-particle density is about 1/5 of human in the same corpus study.

Two additions that may be of general interest

Both came out of baseline testing (running "de-AI-ify this" on Chinese AI text without a skill, then reading what went wrong), so they may generalize beyond Chinese:

  1. A register gate. For encyclopedic, technical, legal, and reference text, neutral plain prose is the correct human voice. Without this, "humanizing" turns a reference article into a personal blog post.
  2. An anti-overcorrection list. The most common failure was not leftover AI tone — it was inventing first-person anecdotes ("back when I was learning to code…") to manufacture human texture. That is fabrication wearing a human mask, so your no-fabrication rule from v2.9.0 is extended to name it explicitly. A second failure was flattening every genre into the same fake-colloquial voice, which is just a different template.

Distinct from #194 (zh-TW): different script, different punctuation standard (“” vs 「」), and different vocabulary; the two are siblings rather than duplicates.

Happy to close this if you'd rather keep the tracker clear — it's here mainly so Simplified Chinese users can find a version that doesn't break their punctuation.

Hi @blader — following #163 (French), #138 (Spanish) and #194 (Traditional Chinese), I built a **Simplified Chinese (zh-CN)** localization of humanizer. Flagging it here for discoverability, not to request a merge. Repo: https://github.com/jiji262/humanizer-chinese You noted in #163 that you'd rather localized variants stay as separate community repositories so each language can evolve without adding duplicate runtime authorities to the core project. This one fits that shape: standalone repo, MIT, with blader/humanizer attribution kept in the LICENSE, README, and SKILL.md. ### Why it is an adaptation, not a translation Three of your rules invert or soften in Simplified Chinese, so a literal port would actively damage correct prose: - **Curly quotes (§19) invert.** `“”` are the *standard* in Simplified Chinese (GB/T 15834-2011), not a tell. The rule is replaced with detection of full-width/half-width punctuation mixing, which is the real machine fingerprint. - **The em dash ban (§14) is softened, not copied.** The Chinese dash `——` is legitimate punctuation for explanation and topic shifts. Only overuse, paired parenthetical dashes, and mis-typesetting (single-width `—`, or `-` / `--` standing in for a dash) are flagged. - **Subjectless sentences (§13) are idiomatic Chinese,** not hidden-actor passives. Flagging them would gut normal writing, so that slot instead targets Europeanized `被`-constructions and translationese. - **Hyphenated word pairs (§26) are dropped** — Chinese has no hyphenated compounds. Net: 24 patterns map one-to-one, 8 collapse into 4, 1 is dropped, and 8 Chinese-specific patterns are added, for 36 total. ### Chinese-specific patterns, with sources - **Lecture-note signposting** (`首先…其次…最后`, `值得注意的是`, `综上所述`) — AI conjunction density measured at roughly 3× human in a 6,586-text parallel corpus study (CCL 2023, Beijing Language and Culture University). - **Parallelism and antithesis stacking** — measured at roughly 6× human density in Chinese AI text, the strongest single quantitative tell found. - **Bureaucratic empty verbs** (`对…进行分析`, `作出调整`) — the Europeanized-Chinese problem Yu Kwang-chung criticized in 1987, which LLM training data amplifies. - **Uniformly written register** — AI modal-particle density is about 1/5 of human in the same corpus study. ### Two additions that may be of general interest Both came out of baseline testing (running "de-AI-ify this" on Chinese AI text *without* a skill, then reading what went wrong), so they may generalize beyond Chinese: 1. **A register gate.** For encyclopedic, technical, legal, and reference text, neutral plain prose *is* the correct human voice. Without this, "humanizing" turns a reference article into a personal blog post. 2. **An anti-overcorrection list.** The most common failure was not leftover AI tone — it was **inventing first-person anecdotes** ("back when I was learning to code…") to manufacture human texture. That is fabrication wearing a human mask, so your no-fabrication rule from v2.9.0 is extended to name it explicitly. A second failure was flattening every genre into the same fake-colloquial voice, which is just a different template. Distinct from #194 (zh-TW): different script, different punctuation standard (`“”` vs `「」`), and different vocabulary; the two are siblings rather than duplicates. Happy to close this if you'd rather keep the tracker clear — it's here mainly so Simplified Chinese users can find a version that doesn't break their punctuation.
Sign in to join this conversation.
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
skills/blader-humanizer#203
No description provided.