Improved Word Representation Learning with Sememes

Topic: Word Representation
Dataset: Sogou-T, HowNet

  • 1889 distinct sememes
  • 2.4 average senses for each word
  • 1.6 average sememes for each sense
  • 42.2% of words have multiple senses.

Methodology:
Sememe-Encoded Word Representation Learning(SE WRL)
This framework regards each word sense as a combination of its sememes, and iteratively performs word sense disambiguation according to their contexts and learn representations of sememes, senses and words by extending Skip-gram in word2vec (Mikolov et al., 2013)

Word, sense, sememe

A). Simple Sememe Aggregation Model
For each word, SSA considers all sememes in all senses of the word together, and represents the target word using the average of all its sememe embeddings.
簡(jiǎn)單的sememe聚合模型在低頻詞上能有更好的表現(xiàn),因?yàn)樵趥鹘y(tǒng)skipgram模型中低頻詞不能得到很好的訓(xùn)練,然而在SSA中通過sememe embeddings 低頻詞被解碼為sememe并通過其他詞得到良好的訓(xùn)練。

B). Sememe Attention over Context Model
The SSA Model replaces the target word embedding with the aggregated sememe embeddings to encode sememe information into word representa- tion learning. However, each word in SSA model still has only one single representation in different contexts, which cannot deal with polysemy of most words. It is intuitive that we should construct distinct embeddings for a target word according to specific contexts, with the favor of word sense annotation in HowNet.

attention

每一個(gè)context word擁有一個(gè)attention weight,attention weight 由目標(biāo)詞w和sense向量之間的相關(guān)度,其中sense向量由其組成sememe向量的平均值表示

Sememe Attention over Context Model

C). Sememe Attention over Target Model
The process can also be applied to select appropriate senses for the target word, by taking context words as attention.

attention

對(duì)于Context Model,只有一個(gè)target word用來學(xué)習(xí)context words 的sense權(quán)重;
對(duì)于Target Model,使用多個(gè)context words 來聯(lián)合學(xué)習(xí)target word 的sense 權(quán)重。
因而Target Model能夠產(chǎn)生更好的語義去模糊化結(jié)果,產(chǎn)生更準(zhǔn)確的語義表示。

Sememe Attention over Target Model
最后編輯于
?著作權(quán)歸作者所有,轉(zhuǎn)載或內(nèi)容合作請(qǐng)聯(lián)系作者
【社區(qū)內(nèi)容提示】社區(qū)部分內(nèi)容疑似由AI輔助生成,瀏覽時(shí)請(qǐng)結(jié)合常識(shí)與多方信息審慎甄別。
平臺(tái)聲明:文章內(nèi)容(如有圖片或視頻亦包括在內(nèi))由作者上傳并發(fā)布,文章內(nèi)容僅代表作者本人觀點(diǎn),簡(jiǎn)書系信息發(fā)布平臺(tái),僅提供信息存儲(chǔ)服務(wù)。

相關(guān)閱讀更多精彩內(nèi)容

  • 1吳老師的思維轉(zhuǎn)的很快,不管榮老師講什么都可以反應(yīng)很快的對(duì)答,真心很厲害,而且越來越幽默,讓人相處起來很舒服2莎莎...
    幸運(yùn)的修行閱讀 222評(píng)論 0 0
  • 小黃,29歲,已是一個(gè)快兩歲寶寶的媽了。我倆是校友,雖我和她同歲,而我卻早她一年畢業(yè)。 在大學(xué)里我還不認(rèn)識(shí)小黃,我...
    肖之諾閱讀 797評(píng)論 0 1

友情鏈接更多精彩內(nèi)容