Text tooling in R assumes that words are separated by whitespace, and Chinese, Japanese and Korean writing does not oblige.
tidytext
splits on whitespace. Chinese and Japanese don’t use whitespace. Give
tidytext an English sentence and you get words; give it
我今天很開心 and you get either one undifferentiated blob or six
isolated characters. Neither is right — 開心 (“happy”) is one word of
two characters, and splitting it destroys the signal you were trying to
measure.
tidycjk tells you what script and language a text column is in, how much of it is CJK, which characters it contains, how wide it will print, and normalises the width variants away — as tidy verbs that return tibbles.
# install.packages("pak")
pak::pak("PursuitOfDataScience/tidyckj")Three verbs take (data, column) and return a tibble, so
they drop into a dplyr pipeline. Two
of them summarise a text column; cjk_tokens() is under
Segmentation below.
library(tidycjk)
posts <- data.frame(
id = 1:5,
text = c(
"我今天很開心", # Chinese
"こんにちは、元気ですか", # Japanese, with kana
"안녕하세요", # Korean
"東京都", # Japanese written without kana
"no CJK here at all"
)
)
cjk_summary(posts, text)
#> # A tibble: 1 × 4
#> n_docs n_with_cjk prop_with_cjk mean_ratio
#> <int> <int> <dbl> <dbl>
#> 1 5 4 0.8 0.8prop_with_cjk answers “how many of these rows are CJK at
all”; mean_ratio answers “how CJK are they”. A corpus of
English with one stray ideograph per row scores high on the first and
near zero on the second.
cjk_char_counts(posts, text)
#> # A tibble: 25 × 5
#> char codepoint script block n
#> <chr> <int> <chr> <chr> <int>
#> 1 我 25105 han CJK Unified Ideographs 1
#> 2 今 20170 han CJK Unified Ideographs 1
#> 3 天 22825 han CJK Unified Ideographs 1
#> 4 很 24456 han CJK Unified Ideographs 1
#> 5 開 38283 han CJK Unified Ideographs 1
#> 6 心 24515 han CJK Unified Ideographs 1
#> 7 こ 12371 hiragana Hiragana 1
#> 8 ん 12435 hiragana Hiragana 1
#> 9 に 12395 hiragana Hiragana 1
#> 10 ち 12385 hiragana Hiragana 1
#> # ℹ 15 more rowscjk_tokens() is the CJK-aware counterpart of tidytext’s
unnest_tokens(). Where a word ends is a fact about a
language rather than about Unicode, so no segmenter is bundled and
engine is required — quietly handing back character tokens
to someone who asked for words is the mistake this package exists to
avoid.
cjk_tokens(posts[1:2, ], text, engine = "character")
#> # A tibble: 17 × 3
#> id text token
#> <int> <chr> <chr>
#> 1 1 我今天很開心 我
#> 2 1 我今天很開心 今
#> 3 1 我今天很開心 天
#> 4 1 我今天很開心 很
#> 5 1 我今天很開心 開
#> 6 1 我今天很開心 心
#> 7 2 こんにちは、元気ですか こ
#> 8 2 こんにちは、元気ですか ん
#> 9 2 こんにちは、元気ですか に
#> 10 2 こんにちは、元気ですか ち
#> 11 2 こんにちは、元気ですか は
#> 12 2 こんにちは、元気ですか 、
#> 13 2 こんにちは、元気ですか 元
#> 14 2 こんにちは、元気ですか 気
#> 15 2 こんにちは、元気ですか で
#> 16 2 こんにちは、元気ですか す
#> 17 2 こんにちは、元気ですか かcjk_segment("hello 中文 world", engine = "character")
#> [[1]]
#> [1] "hello" "中" "文" "world""character" is the one engine that needs no dictionary:
one token per CJK character, non-CJK runs split on whitespace. It is
character tokenisation, not word segmentation.
For real Chinese word segmentation, register jiebaR. It was
archived from CRAN on 2025-05-01, so install it with
remotes::install_github("qinwf/jiebaR"), then:
register_cjk_segmenter("jiebar", function(x, ...) {
worker <- jiebaR::worker(...)
lapply(x, function(s) {
if (is.na(s)) return(NA_character_)
if (!nzchar(s)) return(character(0))
as.character(jiebaR::segment(s, worker))
})
})
cjk_segment("我今天很開心", engine = "jiebar")The same shape works for any segmenter you can call from R.
cjk_script(posts$text)
#> [1] "han" "hiragana" "hangul" "han" NA
cjk_detect_language(posts$text)
#> [1] NA "japanese" "korean" NA NARow 4 is NA, and that is the interesting one. 東京都 is
Tokyo Metropolis — Japanese — but it is written entirely in Han
characters, and nothing in the script distinguishes it from Chinese. Any
library that answers "chinese" there is guessing, and it
will be wrong on Japanese input in a way you cannot detect downstream.
tidycjk returns NA.
If your corpus is known to be Chinese, ask for the guess explicitly, so that the assumption is visible in the script rather than buried in a default:
cjk_detect_language(posts$text, han_only = "chinese")
#> [1] "chinese" "japanese" "korean" "chinese" NAnchar() counts characters. A terminal counts columns,
and a CJK character takes two of them — which is why every console table
containing CJK text comes out ragged.
labels <- c("中文", "abcd", "日本語")
nchar(labels) # all look comparable
#> [1] 2 4 3
cjk_width(labels) # they are not
#> [1] 4 4 6Pad and truncate by width and the columns line up:
cat(paste0("|", cjk_pad(labels, 8), "|"), sep = "\n")
#> |中文 |
#> |abcd |
#> |日本語 |
cjk_truncate("我今天很開心", 8)
#> [1] "我今..."The usual advice is to run NFKC. NFKC does
fix width, and it also rewrites ligatures, Roman numerals, circled
numbers and compatibility ideographs. to_halfwidth()
changes width and nothing else.
to_halfwidth("123") # fullwidth digits that would not parse
#> [1] "123"
as.numeric(to_halfwidth("123"))
#> [1] 123
to_halfwidth("½ Ⅸ ①") # NFKC would mangle all three; this leaves them
#> [1] "½ Ⅸ ①"Halfwidth katakana writes a voiced syllable as two code points. Composing them into one is the difference between text that matches a literal and text that silently doesn’t:
x <- "ガ" # U+FF76 U+FF9E — two code points
nchar(x)
#> [1] 2
nchar(to_halfwidth(x)) # one: U+30AC
#> [1] 1
to_halfwidth(x) == "ガ"
#> [1] TRUEEverything above has a stringr-style
counterpart on plain character vectors: has_cjk(),
cjk_script(), cjk_detect_language(),
cjk_ratio(), cjk_width(),
cjk_pad(), cjk_truncate(),
to_halfwidth() and to_fullwidth(). All are
vectorised, propagate NA, and return zero-length output for
zero-length input.
cjk_blocks() exports the block table the whole package
is built on, so the definition of “CJK” can be read rather than guessed
at.
head(cjk_blocks(), 8)
#> # A tibble: 8 × 5
#> block script start end n_codepoints
#> <chr> <chr> <int> <int> <int>
#> 1 Hangul Jamo hangul 4352 4607 256
#> 2 CJK Symbols and Punctuation punctuation 12288 12351 64
#> 3 Hiragana hiragana 12352 12447 96
#> 4 Katakana katakana 12448 12543 96
#> 5 Bopomofo bopomofo 12544 12591 48
#> 6 Hangul Compatibility Jamo hangul 12592 12687 96
#> 7 Kanbun kanbun 12688 12703 16
#> 8 Bopomofo Extended bopomofo 12704 12735 32| Need | Use |
|---|---|
| Display width, width-aware padding | stringi —
stri_width() and stri_pad(), which
cjk_width() and cjk_pad() wrap |
| Chinese segmentation | jiebaR —
archived from CRAN 2025-05-01; register it as a tidycjk
engine |
| Pinyin | pinyin, hanyupinyin |
| Traditional/simplified conversion | tmcn; OpenCC outside R for phrase-level accuracy |
| Japanese utilities | zipangu, Nippon |
GPL (>= 3).