如何解决R按组提取列中最常用的单词/ ngram
我希望从“标题”列中为每个组(第一列)提取主要关键字。
“所需标题”列中的所需结果:
可复制的数据:
myData <-
structure(list(group = c(1,1,2,3,3),title = c("mentoring aug 8th 2018","mentoring aug 9th 2017","mentoring aug 9th 2018","mentoring august 31","mentoring blue care","mentoring cara casual","mentoring CDP","mentoring cell douglas","mentoring centurion","mentoring CESO","mentoring charlotte","medication safety focus","medication safety focus month","medication safety for nurses 2017","medication safety formulations errors","medication safety foundations care","medication safety general","communication surgical safety","communication tips","communication tips for nurses","communication under fire","communication webinar","communication welling","communication wellness")),row.names = c(NA,-24L),class = c("tbl_df","tbl","data.frame"))
我研究了唱片链接解决方案,但这主要是为了对完整标题进行分组。 任何建议都会很棒。
解决方法
我按组将所有标题串联起来,并将它们标记化:
library(dplyr)
myData <-
topic_modelling %>%
group_by(group) %>%
mutate(titles = paste0(title,collapse = " ")) %>%
select(group,titles) %>%
distinct()
myTokens <- myData %>%
unnest_tokens(word,titles) %>%
anti_join(stop_words,by = "word")
myTokens
# finding top ngrams
library(textrank)
stats <- textrank_keywords(myTokens$word,ngram_max = 3,sep = " ")
stats <- subset(stats$keywords,ngram > 0 & freq >= 3)
head(stats,5)
在将算法应用于大约100000行的真实数据时,我创建了一个函数来逐组解决问题:
# FUNCTION: TOP NGRAMS ----
find_top_ngrams <- function(titles_concatenated)
{
myTest <-
titles_concatenated %>%
as_tibble() %>%
unnest_tokens(word,value) %>%
anti_join(stop_words,by = "word")
stats <- textrank_keywords(myTest$word,ngram_max = 4,sep = " ")
stats <- subset(stats$keywords,ngram > 1 & freq >= 5)
top_ngrams <- head(stats,5)
top_ngrams <- tibble(top_ngrams)
return(top_ngrams)
# print(top_ngrams)
}
for (i in 1:5){
find_top_ngrams(myData$titles[i])
}
版权声明:本文内容由互联网用户自发贡献,该文观点与技术仅代表作者本人。本站仅提供信息存储空间服务,不拥有所有权,不承担相关法律责任。如发现本站有涉嫌侵权/违法违规的内容, 请发送邮件至 dio@foxmail.com 举报,一经查实,本站将立刻删除。