ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

从零训练自定义词向量:在 nlp-recipes 中实战 Word2Vec、GloVe 与 fastText

从零训练自定义词向量:在 nlp-recipes 中实战 Word2Vec、GloVe 与 fastText NLP【免费下载链接】nlp-recipesNatural Language Processing Best Practices Examples项目地址https://gitcode.com/gh_mirrors/nl/nlp-recipes点击查看免费下载导读本文基于 nlp-recipes 仓库中 examples/embeddings 目录的官方示例系统讲解如何用 Word2Vec、GloVe 与 fastText 三种经典方法在自定义语料上从零训练词向量Word Embedding。当通用预训练模型如基于 Wikipedia、Common Crawl 训练的词向量无法覆盖领域专属语言或小语种场景时这套流程可以直接落地。读完本文你将掌握数据预处理工具链的调用方式、gensim 训练 Word2Vec/fastText 的关键参数、GloVe 四步 C 程序流水线以及如何保存、加载与检查训练产出的词向量。为什么要自己训练词向量词向量Word Embedding是把词汇表vocabulary中的词或短语映射为实数向量的技术学习到的向量表示能够捕获词之间的句法关系与语义关系因此对句子相似度sentence similarity、文本分类text classification等下游任务非常有用。社区中已有大量开箱即用的预训练模型Word2Vec、GloVe、fastText 三种方法均发布了基于 Wikipedia、Common Crawl 等通用语料训练的公开版本。但通用模型在两类场景下并不适用领域专属语言问题例如医学、法律、代码等垂直领域的行话与专有表达通用语料中的统计分布无法覆盖无预训练模型的语言某些小语种没有公开可用的预训练词向量。此时利用自己的语料从零训练词向量是标准解法。本仓库在 examples/embeddings/embedding_trainer.ipynb 中给出了完整可运行的 Jupyter Notebook 示例覆盖数据加载、预处理、三种模型训练与结果检查的完整链路数据集选用 STS Benchmark英文。示例速览Notebook环境说明数据集语言Developing Word Embeddings本地演示用 Word2Vec、fastText 与 GloVe 学习词表示STS Benchmark dataseten环境准备与路径规划Notebook 的第一步是导入依赖并规划目录结构。训练词向量主要依赖gensim提供 Word2Vec 与 FastText 实现以及仓库自带的工具模块import gensim import sys import os # Set the environment path sys.path.append(../..) import numpy as np from utils_nlp.dataset.preprocess import ( to_lowercase, to_spacy_tokens, rm_spacy_stopwords, ) from utils_nlp.dataset import stsbenchmark from utils_nlp.common.timer import Timer from gensim.models import Word2Vec from gensim.models.fasttext import FastText路径配置上Notebook 定义了三个关键目录# Set the path for where your repo is located NLP_REPO_PATH os.path.join(..,..) # Set the path for where your datasets are located BASE_DATA_PATH os.path.join(NLP_REPO_PATH, data) # Set the path for location to save embeddings SAVE_FILES_PATH os.path.join(BASE_DATA_PATH, trained_word_embeddings) if not os.path.exists(SAVE_FILES_PATH): os.makedirs(SAVE_FILES_PATH)NLP_REPO_PATH仓库根目录用于定位utils_nlp包BASE_DATA_PATH数据集存放目录原始数据会被下载到data/raw下SAVE_FILES_PATH所有训练产出的词向量与中间文件语料、词汇表、共现矩阵统一保存到data/trained_word_embeddings。仓库中的底层支撑模块utils_nlp/common/timer.py提供Timer类封装timeit.default_timer通过start()/stop()或with Timer() as t两种用法计时print(t)直接输出秒数保留 4 位小数用于对比三种模型的训练耗时utils_nlp/dataset/stsbenchmark.py负责 STS Benchmark 的下载、解压与 DataFrame 化utils_nlp/dataset/preprocess.py提供小写化、spaCy 分词、停用词过滤等预处理函数。数据加载与预处理加载 STS Benchmark使用仓库封装好的stsbenchmark模块加载训练集并清洗掉无关元数据列# Produce a pandas dataframe for the training set train_raw stsbenchmark.load_pandas_df(BASE_DATA_PATH, file_splittrain) # Clean the sts dataset sts_train stsbenchmark.clean_sts(train_raw)stsbenchmark.load_pandas_df会从官方地址下载Stsbenchmark.tar.gz并解压到BASE_DATA_PATH/raw下然后读取sts-train.csv通过file_split参数还可加载dev与test分片。清洗后的 DataFrame 只保留三列scoresentence1sentence25.00A plane is taking off.An air plane is taking off.3.80A man is playing a large flute.A man is playing a flute.训练集清洗后规模为5749 行 × 3 列sts_train.shape的输出即 5749 对带相似度打分的句子。底层实现细节clean_sts会丢弃原始的column_0column_3体裁、标注来源等元数据仅保留score、sentence1、sentence2三列见 utils_nlp/dataset/stsbenchmark.py。训练集预处理对句子对执行三步标准预处理全部来自 utils_nlp/dataset/preprocess.py# Convert all text to lowercase df_low to_lowercase(sts_train) # Tokenize text sts_tokenize to_spacy_tokens(df_low) # Tokenize with removal of stopwords sts_train_stop rm_spacy_stopwords(sts_tokenize)to_lowercase(df, column_names[])将所有字符串列转为小写不传列名时作用于整个 DataFrameto_spacy_tokens(df)用 spaCy 的en_core_web_sm模型把sentence1、sentence2分词产出sentence1_tokens、sentence2_tokens两列每列元素为 token 列表rm_spacy_stopwords(df)在上一步 token 列基础上过滤掉停用词产出sentence1_tokens_rm_stopwords、sentence2_tokens_rm_stopwords该函数还支持通过custom_stopwords参数向 spaCy 模型注册自定义停用词nlp.vocab[csw].is_stop True适合领域术语过滤。随后把两列 token 拼接展平得到模型训练所需的句子列表并过滤掉空句all_sentences sts_train_stop[[sentence1_tokens_rm_stopwords, sentence2_tokens_rm_stopwords]] # Flatten two columns into one list and remove all sentences that are size 0 sentences [i for i in all_sentences.values.flatten().tolist() if len(i) 0]预处理后共得到11498 条句子len(sentences)句子长度统计为最短 1 个 token、最长 43 个 token、中位数 6 个 token。抽样查看前 10 条可以看到停用词已被移除[[plane, taking, .], [air, plane, taking, .], [man, playing, large, flute, .], ...]方法一Word2Vec原理Word2Vec 是一种**预测式predictive**词向量学习方法。其核心思想是在向量空间中语料中共享上下文context的词彼此靠近。它有两种经典模型架构CBOWcontinuous bag-of-words用窗口内周围的词上下文预测当前词Skip-gram用当前词预测周围的上下文词。关键参数gensim 的Word2Vec参数很多Notebook 中特别标注了最常用的五个参数含义默认值size词向量维度100window被预测词与当前词之间的最大距离上下文窗口大小5min_count忽略出现频率低于该值的所有词5workers训练使用的 worker 线程数3sg训练算法1 为 skip-gram0 为 CBOW0训练与计时t Timer() t.start() # Train the Word2vec model word2vec_model Word2Vec(sentences, size100, window5, min_count5, workers3, sg0) t.stop() print(Time elapsed: {}.format(t)) # Time elapsed: 0.3194在 11498 条句子的小语料上Word2Vec 模型训练耗时约0.32 秒Notebook 中训练单元的运行结果。检查与保存训练完成后可以完成三件事# 1. 查询某个词的词向量通过 wv 属性以词为 key 访问 print(Embedding for apple:, word2vec_model.wv[apple]) # 2. 查看模型词汇表wv.vocab 的 key print(\nFirst 30 vocabulary words:, list(word2vec_model.wv.vocab)[:20]) # 3. 保存词向量二进制格式节省空间或 ASCII 格式 word2vec_model.wv.save_word2vec_format(SAVE_FILES_PATHword2vec_model, binaryTrue) # binary word2vec_model.wv.save_word2vec_format(SAVE_FILES_PATHword2vec_model, binaryFalse) # ASCIINotebook 输出显示apple的词向量是 100 维浮点数组词汇表前 20 个词为[plane, taking, ., air, man, playing, large, flute, spreading, cheese, pizza, men, seated, fighting, smoking, piano, guitar, singing, woman, person]——可见训练语料主题集中在图片描述类句子。save_word2vec_format输出的文本/二进制格式可被gensim.models.KeyedVectors.load_word2vec_format重新加载。方法二fastText原理fastText 是 Facebook Research 提出的无监督词向量学习算法。它与 Word2Vec、GloVe 的本质区别在于最小单元Word2Vec 与 GloVe 把每个词当作不可再分的最小单元fastText 假设词由字符 n-gram组成。例如单词 language 的 2-gram 为 {la, an, ng, gu, ua, ag, ge}词的嵌入由这些字符 n-gram 的向量求和得到。这一设计带来的直接收益是稀有词与词表外OOV词仍可被拆解为字符 n-gram 而得到向量因此在小数据集上 fastText 通常比 Word2Vec 和 GloVe 表现更好。关键参数gensim 的FastText参数与 Word2Vec 基本一致另多出iter参数含义默认值size词向量维度100window上下文窗口大小5min_count忽略频率低于该值的词5workersworker 线程数3sg1 为 skip-gram0 为 CBOW0iter训练轮数epochs5训练与检查t.start() # Train the FastText model fastText_model FastText(size100, window5, min_count5, sentencessentences, iter5) t.stop() print(Time elapsed: {}.format(t)) # Time elapsed: 9.3665由于需要额外学习字符 n-gram 的表示fastText 在同一语料上耗时显著高于 Word2Vec本示例约9.37 秒。由于两者同源于 gensim 包可以使用完全相同的 API 检查与保存print(Embedding for apple:, fastText_model.wv[apple]) print(\nFirst 30 vocabulary words:, list(fastText_model.wv.vocab)[:20]) fastText_model.wv.save_word2vec_format(SAVE_FILES_PATHfastText_model, binaryTrue) # binary fastText_model.wv.save_word2vec_format(SAVE_FILES_PATHfastText_model, binaryFalse) # ASCII注意fastText 的save_word2vec_format保存的是词向量本身若需保留字符 n-gram 信息以支持 OOV 词查询应使用 gensim 的完整模型保存方式model.save()。方法三GloVe原理与工程背景GloVeGlobal Vectors for Word Representation是 Stanford NLP 团队提出的无监督词向量算法。它基于词-词共现统计word-word co-occurrence statistics训练学习目标是让两个词的向量点积等于它们的共现概率从而同时利用全局统计信息与局部上下文信息。关键工程背景gensim 没有实现 GloVe 模型而其他 Python 包实现不稳定因此本仓库直接借用了 Stanford NLP 官方 GloVe 仓库 模块包含src/glove.c、src/vocab_count.c、src/cooccur.c、src/shuffle.c四个 C 源文件与 Makefile。训练前需要先编译 C 程序# Define path glove_model_path os.path.join(NLP_REPO_PATH, utils_nlp, models, glove) # Execute shell commands !cd $glove_model_path makemake会根据 Makefile 依次编译出build/glove、build/shuffle、build/cooccur、build/vocab_count四个可执行文件编译参数为-lm -pthread -Ofast -marchnative -funroll-loops。编译过程中会出现fread返回值未检查的 warning属于官方源码的已知现象不影响使用。Step 0准备数据GloVe 的 C 工具链要求语料为纯文本文件所有词以 1 个以上空格或 Tab 分隔每个文档/句子以换行符分隔。将前面预处理得到的句子列表写入文件training_corpus_file_path os.path.join(SAVE_FILES_PATH, training-corpus-cleaned.txt) with open(training_corpus_file_path, w, encodingutf8) as file: for sent in sentences: file.write( .join(sent) \n)Step 1构建词汇表vocab_count运行vocab_count可执行文件统计词频并构建词汇表可选参数参数含义min-count词在数据集中出现次数的下限低于该值的词从词汇表丢弃max-vocab保留词汇词数的上限verbose日志级别0、1 或 2默认vocab_count_exe_path os.path.join(glove_model_path, build, vocab_count) vocab_file_path os.path.join(SAVE_FILES_PATH, vocab.txt) !$vocab_count_exe_path -min-count 5 -verbose 2 $training_corpus_file_path $vocab_file_pathNotebook 的运行输出展示了实际统计结果共处理 85334 个 token统计到 11716 个唯一词按min count 5截断后词汇表规模为2943个词。Step 2构建共现统计cooccur运行cooccur统计词对共现主要参数参数含义symmetric0 只考虑左上下文1默认同时考虑左右上下文window-size使用的上下文词数量默认 15verbose0、1 或 2默认vocab-fileStep 1 生成的词汇表文件路径memory内存消耗软限制默认 4单位 GBmax-product通过限定两个共现词频次乘积的最大整数值控制稠密共现数组大小cooccur_exe_path os.path.join(glove_model_path, build, cooccur) cooccurrence_file_path os.path.join(SAVE_FILES_PATH, cooccurrence.bin) !$cooccur_exe_path -memory 4 -vocab-file $vocab_file_path -verbose 2 -window-size 15 $training_corpus_file_path $cooccurrence_file_pathNotebook 输出窗口大小 15、对称上下文、读取 2943 个词的词汇表、处理 85334 个 token、合并共现文件共 188154 行。Step 3打乱共现数据shuffle运行shuffle打乱共现记录参数如下参数含义verbose0、1 或 2默认memory内存消耗软限制默认 4array-size写盘前缓冲区的最大长度限制每次打乱的数据块大小shuffle_exe_path os.path.join(glove_model_path, build, shuffle) cooccurrence_shuf_file_path os.path.join(SAVE_FILES_PATH, cooccurrence.shuf.bin) !$shuffle_exe_path -memory 4 -verbose 2 $cooccurrence_file_path $cooccurrence_shuf_file_path打乱共现数据是 SGD 训练前的必要步骤目的是避免相邻共现记录之间的相关性干扰梯度更新。Step 4训练 GloVe 模型glove运行glove主程序训练模型常用参数参数含义verbose0、1 或 2默认vector-size词向量维度默认 50threads线程数默认 8iter迭代次数默认 25eta学习率默认 0.05binary保存格式0 为文本默认、1 为二进制、2 为两者都保存x-max加权函数的截断阈值默认 100vocab-fileStep 1 生成的词汇表文件save-file向量保存文件名input-fileStep 3 产出的共现文件glove_exe_path os.path.join(glove_model_path, build, glove) glove_vector_file_path os.path.join(SAVE_FILES_PATH, GloVe_vectors) !$glove_exe_path -save-file $glove_vector_file_path -threads 8 -input-file \ $cooccurrence_shuf_file_path -x-max 10 -iter 15 -vector-size 50 -binary 2 \ -vocab-file $vocab_file_path -verbose 2Notebook 训练日志可见关键配置向量维度 50、词汇表 2943、x_max: 10、alpha: 0.75加权函数指数共迭代 15 轮cost 从第 1 轮的 0.078545 逐步下降到第 15 轮的 0.041065。整个 GloVe 流水线Step 04在示例中耗时约3.43 秒。检查 GloVe 词向量GloVe 输出的是纯文本向量文件binary2时同时生成.txt与.bin两种格式。Notebook 手动读取文本文件构建词向量字典glove_wv {} glove_vector_txt_file_path os.path.join(SAVE_FILES_PATH, GloVe_vectors.txt) with open(glove_vector_txt_file_path, encodingutf-8) as f: for line in f: split_line line.split( ) glove_wv[split_line[0]] [float(i) for i in split_line[1:]]与 gensim 模型类似可以查询apple的 50 维向量并查看词汇表前 20 个词[., ,, man, -, , woman, , said, dog, playing, :, white, black, $, killed, percent, new, syria, people, china]。注意 GloVe 词汇表与 Word2Vec/fastText 不同——它基于全量语料统计而非仅训练词向量且默认保留标点符号等 token。三种方法的对比与选型建议Notebook 的 Concluding Remarks 汇总了三种方法在本示例语料上的整体耗时Word2Vec 约 0.39 秒、GloVe 约 8.16 秒、fastText 约 10.41 秒这里的 GloVe 与 fastText 耗时包含完整训练流程。方法学习方式最小单元OOV 词支持小数据集表现本示例整体耗时Word2Vec预测式CBOW/Skip-gram词否一般~0.39 sGloVe基于全局共现统计词否一般~8.16 sfastText预测式 字符 n-gram字符 n-gram是通常更好~10.41 s选型建议来自 Notebook 的结论fastText 通常被视为词向量的良好基线适合作为生成词向量的起点其对 OOV 词和稀有词的鲁棒性使其尤其适合词汇覆盖不足的领域语料。若数据量极大且追求训练速度Word2VecCBOW的预测式训练更具效率优势。训练后的进阶用法训练好的词向量并非终点。Notebook 明确指出生成自训练词向量后可以复用仓库中 examples/sentence_similarity/baseline_deep_dive.ipynb 的句子相似度流程——将其中使用的互联网预训练词向量替换成本文训练出的向量即可在同一套框架下评估自训练向量的质量。此外仓库还提供了预训练向量的加载封装可用于与自训练向量做对比实验utils_nlp/models/pretrained_embeddings/word2vec.py加载 GoogleNews 预训练向量utils_nlp/models/pretrained_embeddings/glove.py加载 GloVe 预训练向量utils_nlp/models/pretrained_embeddings/fasttext.py加载 fastText 预训练向量。对应的冒烟测试 tests/smoke/test_word_embeddings.py 验证了三种加载函数均能把词向量载入为gensim.models.keyedvectors.Word2VecKeyedVectors或FastText对象并检查了下载文件大小与词汇规模——这说明仓库对预训练向量加载与自训练向量两条路径都提供了统一的数据结构支撑训练产出可以无缝接入后续下游任务。结语通过 examples/embeddings/embedding_trainer.ipynb 的完整示例可以快速掌握在自定义语料上训练词向量的三种主流方案gensim 一行 API 即可完成 Word2Vec 与 fastText 训练而 GloVe 需要走语料文本 → vocab_count 建词表 → cooccur 统计共现 → shuffle 打乱 → glove 训练的四步 C 工具链。无论选择哪种方法核心都归结为三点预处理质量决定输入、超参数决定形态、语料规模决定上限。当通用预训练模型无法覆盖你的领域词汇或目标语言时这套自训练流程就是最直接的替代方案。赞分享NLP【免费下载链接】nlp-recipesNatural Language Processing Best Practices Examples项目地址https://gitcode.com/gh_mirrors/nl/nlp-recipes点击查看免费下载相关推荐使用 nlp-recipes 加载预训练词向量fastText / GloVe / Word2Vec 一键下载与加载实战使用 nlp recipes 加载预训练词向量fastText / GloVe / Word2Vec 一键下载与加载实战 导读 nlp recipes 的 uNLP预训练词向量终极指南Word2Vec、GloVe、FastText深度对比预训练词向量终极指南Word2Vec、GloVe、FastText深度对比 在自然语言处理NLP领域预训练词向量是构建智能文本应用的基础技术。nlp rNLPnlp-recipes 中的 GloVe 训练工具链从语料预处理到词向量训练的完整实战指南nlp recipes 中的 GloVe 训练工具链从语料预处理到词向量训练的完整实战指南 导读 本篇文章以 nlp recipes 仓库内 utils_nlNLP上一篇终极Anima-LLLite姿势控制指南10个技巧实现AI人物姿态精准生成下一篇gorp vs 原生SQL为什么这个ORM-ish库能让你的代码更优雅创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表