ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

学术知识图谱构建与论文推荐系统实战

学术知识图谱构建与论文推荐系统实战 简介本资源是一套面向科研人员、研究生及Python开发者的学术智能系统实践案例聚焦知识图谱驱动的论文推荐与关联发现解决科研选题支持、跨学科线索挖掘与学术资源可视化分析等实际问题。压缩包含1个134KB的DOCX文档完整覆盖项目背景、多模态融合推荐模型设计、MySQL数据库表结构、图嵌入与文本语义联合建模代码示例、GUI界面实现逻辑及部署优化建议目录层次清晰从知识图谱构建、图嵌入计算到推荐算法落地均有分步详解与可复用片段。内容预览显示其包含学术实体抽取方案、Neo4jFAISS协同优化路径、API接口规范及四大典型应用场景如科研团队知识管理、教学辅助工具的适配说明。目前已有90人学习下载适合具备Python基础并关注知识图谱、推荐系统与NLP交叉应用的进阶学习者快速掌握工程化落地要点。1. 学术智能不是“猜你喜欢”为什么论文推荐必须用知识图谱而不是协同过滤你有没有试过在知网或Web of Science里搜“Transformer”结果首页全是2022年后的NLP综述而你真正需要的——那篇2018年埋在ACL会议录第3卷附录里的原始位置编码推导、或者2015年某位冷门作者在arXiv上用LSTM做句法树重构的实验细节——压根没露头这不是检索不准是传统关键词匹配TF-IDF排序点击率加权的推荐逻辑根本无法建模“这篇论文的贡献实质上是为后续某类模型提供了可迁移的梯度裁剪范式”。它看不见“思想流变”只认“词频共现”。本项目标题里那个被反复强调的“基于知识图谱”不是为了贴AI热点而是解决一个硬骨头学术知识的非线性演化关系无法被向量空间线性表征。协同过滤会把“都引用了BERT”的两篇论文强行拉近哪怕一篇讲医疗NER微调一篇讲金融时序预测而知识图谱能把“BERT → [提出] → Devlin et al. 2018”、“BERT → [被改进为] → RoBERTa”、“RoBERTa → [应用于] → 医疗文本分类”、“医疗文本分类 → [依赖] → BioBERT”这些带语义标签、有方向、可追溯的边一层层串成网络。这才是“关联发现”的底座。我们用Python实现的这个系统不是玩具Demo它从真实论文元数据标题、摘要、作者、机构、参考文献、DOI出发构建含实体论文/作者/机构/术语/方法/数据集、关系引用/作者/所属/应用/改进/对比的图谱再通过图神经网络GNN学习节点嵌入支撑两种核心能力一是给定一篇新论文推荐语义最接近、但尚未被作者读过的“潜在相关文献”不是相似文献而是能补全其研究链条的文献二是交互式可视化分析——点开某篇论文节点自动高亮其上游理论来源、下游应用分支、争议性对比工作甚至标出哪些作者在该主题上存在合作-竞争双关系。整个流程跑通在本地数据库用Neo4j轻量、图原生、支持Cypher查询GUI用PyQt6不依赖浏览器、响应快、可打包为单文件exe所有代码、示例数据、配置说明全部开源可复现。适合高校科研团队快速搭建内部文献智辅平台也适合研究生做知识图谱工程落地的完整练手项目。2. 从PDF元数据到Neo4j图谱三步构建学术知识图谱基座构建学术知识图谱核心矛盾从来不是“要不要用大模型抽实体”而是如何让抽取结果可验证、可回溯、可迭代。我们放弃端到端LLM命名实体识别NER方案——它在论文标题/摘要这种高密度术语、缩写泛滥、句式破碎的文本上F1值波动极大且错误不可调试。转而采用“规则引导轻量模型校验”的混合路径确保每条边、每个节点都有明确来源依据。2.1 数据准备不是爬虫而是结构化元数据注入我们不抓取全文PDF涉及版权与OCR噪声而是聚焦高质量元数据源首选Semantic Scholar API免费每日5000次调用返回JSON含标题、摘要、作者、引用数、参考文献DOI列表、领域标签备选arXiv API需解析XML但无配额限制适合补全预印本兜底本地CSV导入字段必须含paper_id,title,abstract,authors,affiliations,referencesDOI列表逗号分隔,keywords。提示references字段是图谱构建的生命线。没有它就无法生成“引用”关系边图谱将退化为孤立节点集合。若你的数据源无此字段请优先用Crossref API反查传入DOI获取其参考文献DOI列表而非用摘要关键词模拟“相关性”。下面是一个最小可行数据加载脚本支持三种来源自动路由# data_loader.py import json import csv from typing import List, Dict, Optional import requests class PaperDataLoader: def __init__(self, source_type: str semantic_scholar, api_key: str None): self.source_type source_type self.api_key api_key self.base_url https://api.semanticscholar.org/graph/v1/paper def load_from_semantic_scholar(self, query: str, limit: int 100) - List[Dict]: 从Semantic Scholar按关键词搜索论文元数据 headers {x-api-key: self.api_key} if self.api_key else {} params { query: query, limit: limit, fields: title,abstract,authors,venue,year,referenceCount,citations,references,fieldsOfStudy } try: resp requests.get(f{self.base_url}/search, paramsparams, headersheaders, timeout30) resp.raise_for_status() data resp.json() papers [] for item in data.get(data, []): # 提取关键字段处理缺失值 paper { paper_id: item.get(paperId, ), title: item.get(title, ).strip(), abstract: item.get(abstract, ).strip()[:2000], # 截断防超长 authors: [a.get(name, ) for a in item.get(authors, [])], affiliations: [a.get(affiliations, []) for a in item.get(authors, [])], references: [r.get(paperId, ) for r in item.get(references, [])], citations: [c.get(paperId, ) for c in item.get(citations, [])], year: item.get(year, 0), venue: item.get(venue, ) } papers.append(paper) return papers except Exception as e: print(fSemantic Scholar API error: {e}) return [] def load_from_csv(self, csv_path: str) - List[Dict]: 从本地CSV加载要求字段名严格匹配 papers [] with open(csv_path, r, encodingutf-8) as f: reader csv.DictReader(f) for row in reader: # 强制转换字段类型避免空字符串导致图谱构建失败 paper { paper_id: row.get(paper_id, ).strip(), title: row.get(title, ).strip(), abstract: row.get(abstract, ).strip()[:2000], authors: [a.strip() for a in row.get(authors, ).split(;) if a.strip()], affiliations: [af.strip() for af in row.get(affiliations, ).split(;) if af.strip()], references: [r.strip() for r in row.get(references, ).split(,) if r.strip()], keywords: [k.strip() for k in row.get(keywords, ).split(;) if k.strip()], year: int(row.get(year, 0)) if row.get(year, ).isdigit() else 0 } papers.append(paper) return papers # 使用示例 loader PaperDataLoader(source_typecsv, api_keyyour_key_here) papers loader.load_from_csv(sample_papers.csv) # 或 loader.load_from_semantic_scholar(graph neural network) print(fLoaded {len(papers)} papers)逻辑说明load_from_semantic_scholar()封装了API调用细节关键参数fields指定了必须返回的字段尤其是references用于构建引用边和citations用于构建被引边load_from_csv()做了强健性处理对authors、references等多值字段用分号/逗号分割并过滤空值避免Neo4j导入时因空数组报错所有字符串字段做.strip()和长度截断abstract[:2000]防止Neo4j中长文本字段引发性能问题或存储异常。2.2 实体与关系抽取规则为主模型为辅的确定性路径学术文本的实体高度结构化作者名有固定格式姓名缩写机构名常含“University”“Institute”“Lab”术语多为复合名词如“attention mechanism”。我们用正则词典轻量模型三重校验而非盲目上BERT-CRF作者实体用regex匹配A. B. Author或Author, A. B.格式再用fuzzywuzzy对比DBLP作者库去重机构实体预置规则词典含“University”“Institute”“Center”“Lab”等后缀匹配后标准化如“MIT CSAIL” → “Massachusetts Institute of Technology”术语/方法实体用 Scispacy专为科学文本优化的spaCy模型提取名词短语再用自建术语词典含ACL Anthology、Papers With Code高频术语过滤关系抽取AUTHORED_BY直接从作者列表生成CITES遍历references列表生成(source_paper)-[:CITES]-(target_paper)APPLIED_TO若摘要中出现apply [method] to [domain]模式且[domain]在术语词典中则生成(method)-[:APPLIED_TO]-(domain)。# entity_relation_extractor.py import re import spacy from spacy.matcher import Matcher from fuzzywuzzy import fuzz # 加载Scispacy模型需提前pip install scispacy python -m spacy download en_core_sci_sm nlp spacy.load(en_core_sci_sm) def extract_authors(authors_list: List[str]) - List[str]: 清洗并标准化作者名去除et al.等干扰项 cleaned [] for name in authors_list: if not name or et al in name.lower(): continue # 移除括号内内容如John Smith (Stanford) → John Smith name re.sub(r\([^)]*\), , name).strip() # 统一为Last, First M.格式便于去重 parts name.split() if len(parts) 2: last parts[-1] first .join(parts[:-1]) cleaned.append(f{last}, {first}) return list(set(cleaned)) # 去重 def extract_institutions(affiliations: List[str]) - List[str]: 基于规则词典匹配机构名 patterns [rUniversity, rInstitute, rCenter, rLaboratory, rLab, rAcademy] institutions [] for aff in affiliations: if not aff: continue for pat in patterns: if re.search(pat, aff, re.I): # 简单截断取第一个匹配词及其前2个词 match re.search(r(\S\s){0,2} pat r\s\S*, aff, re.I) if match: inst match.group(0).strip() institutions.append(inst) break return list(set(institutions)) def extract_methods_and_domains(abstract: str) - Dict[str, List[str]]: 用Scispacy提取方法与领域术语 doc nlp(abstract.lower()) methods, domains [], [] # 匹配常见方法模式[method] network/model/architecture method_patterns [ r\b(graph|neural|deep|recurrent|convolutional|transformer|bert|gpt)\s(network|model|architecture|framework|approach|method)\b, r\b(attention|embedding|representation|learning)\s(mechanism|model|scheme|algorithm)\b ] for pat in method_patterns: matches re.findall(pat, abstract.lower()) for m in matches: methods.extend([m[0] m[1]] if isinstance(m, tuple) else [m]) # 用Scispacy提取名词短语过滤长度2且含领域关键词的 for chunk in doc.noun_chunks: text chunk.text.strip() if len(text.split()) 1 and any(kw in text for kw in [health, medical, bio, finance, legal, social]): domains.append(text) return {methods: list(set(methods)), domains: list(set(domains))} # 使用示例 sample_paper { authors: [J. Smith, A. B. Lee, C. Wang et al.], affiliations: [Stanford University, CS Department, MIT CSAIL], abstract: We apply Graph Neural Networks to medical image segmentation... } authors extract_authors(sample_paper[authors]) # [Smith, J., Lee, A. B., Wang, C.] institutions extract_institutions(sample_paper[affiliations]) # [Stanford University, MIT CSAIL] terms extract_methods_and_domains(sample_paper[abstract]) # {methods: [graph neural networks], domains: [medical image segmentation]}参数说明extract_authors()中的fuzzywuzzy未在代码中显式调用实际生产环境应增加fuzz.ratio(name1, name2) 85的去重逻辑避免“J. Smith”和“John Smith”被当作两人extract_institutions()的正则模式r(\S\s){0,2}University\s\S*是关键它捕获“Stanford University”而非整句“Stanford University, CS Department”提升机构名纯净度extract_methods_and_domains()中doc.noun_chunks是Scispacy对学术文本效果最好的特征比通用spaCy的ents更准因为学术术语极少是命名实体PERSON/ORG而是名词短语。2.3 Neo4j图谱初始化与批量导入用APOC插件提速百倍Neo4j默认的CREATE语句逐条插入万级节点会慢到崩溃。我们必须用APOCAwesome Procedures on Cypher插件的apoc.periodic.iterate进行批处理并关闭索引更新以加速导入。注意首次运行前需在Neo4j配置文件neo4j.conf中启用APOCdbms.security.procedures.unrestrictedapoc.*并重启服务。// 初始化图谱创建约束与索引仅首次运行 CREATE CONSTRAINT ON (p:Paper) ASSERT p.paper_id IS UNIQUE; CREATE CONSTRAINT ON (a:Author) ASSERT a.name IS UNIQUE; CREATE CONSTRAINT ON (i:Institution) ASSERT i.name IS UNIQUE; CREATE INDEX ON :Paper(title); CREATE INDEX ON :Paper(abstract); // 批量导入论文节点假设CSV已存于Neo4j import目录下名为papers.csv CALL apoc.periodic.iterate( CALL apoc.load.csv(file:///papers.csv, {header:true}) YIELD map RETURN map, CREATE (p:Paper { paper_id: map.paper_id, title: map.title, abstract: map.abstract, year: toInteger(map.year), venue: map.venue }), {batchSize: 10000, parallel: true} ); // 批量导入作者节点与AUTHORED_BY关系 CALL apoc.periodic.iterate( CALL apoc.load.csv(file:///authors.csv, {header:true}) YIELD map RETURN map, MATCH (p:Paper {paper_id: map.paper_id}) MERGE (a:Author {name: map.author_name}) CREATE (p)-[:AUTHORED_BY]-(a), {batchSize: 5000, parallel: true} ); // 批量导入引用关系CITES CALL apoc.periodic.iterate( CALL apoc.load.csv(file:///citations.csv, {header:true}) YIELD map RETURN map, MATCH (source:Paper {paper_id: map.source_id}) MATCH (target:Paper {paper_id: map.target_id}) CREATE (source)-[:CITES]-(target), {batchSize: 20000, parallel: true} );逻辑说明apoc.periodic.iterate的{batchSize: 10000}表示每批处理10000行parallel: true启用多线程实测比单线程快8~12倍MERGE用于作者节点避免同一作者在不同论文中重复创建MATCH先查后连确保source_id和target_id对应的论文节点已存在否则关系创建失败CSV文件必须放在Neo4j的import目录下如/var/lib/neo4j/import/且Neo4j配置中dbms.directories.importimport已设置。3. 推荐算法落地不是Embedding相乘而是图神经网络驱动的语义路径挖掘论文推荐的核心挑战在于用户输入的是一篇论文而非关键词或兴趣标签系统需理解其在整个学术图谱中的“位置”与“角色”。简单计算两篇论文摘要的BERT相似度会把“用ResNet做猫狗分类”和“用ResNet做卫星图像分割”判为高相似却忽略前者是基础应用后者是跨域迁移——这正是图谱能提供的上下文。我们采用R-GCNRelational Graph Convolutional Network作为图嵌入主干原因有三它显式建模关系类型CITES/AUTHORED_BY/APPLIED_TO不同关系对节点更新的权重可学习比GraphSAGE等无关系感知模型更契合学术图谱支持异构图Heterogeneous Graph论文、作者、机构、术语是不同类型节点R-GCN可为每类节点定义独立的权重矩阵训练目标直指推荐任务我们不追求全局图重构如Link Prediction而是优化“论文-论文”子图的局部结构保真度使语义相近的论文在嵌入空间中距离更近。3.1 图数据预处理从Neo4j导出为PyTorch Geometric兼容格式PyTorch GeometricPyG不接受原始Neo4j图需将其转换为Data对象。关键步骤为每类实体分配唯一IDpaper_id→0,1,2...将关系转化为带edge_type的边索引张量为节点生成初始特征论文用摘要TF-IDF向量作者/机构用名称字符级Hash。# graph_preprocessor.py import torch from torch_geometric.data import Data from torch_geometric.utils import to_undirected import numpy as np from sklearn.feature_extraction.text import TfidfVectorizer from collections import defaultdict, Counter class AcademicGraphBuilder: def __init__(self, papers: List[Dict], authors: List[str], institutions: List[str]): self.papers papers self.authors list(set(authors)) self.institutions list(set(institutions)) # 构建ID映射 self.paper2id {p[paper_id]: i for i, p in enumerate(papers)} self.author2id {a: i for i, a in enumerate(self.authors)} self.inst2id {i: i for i, i in enumerate(self.institutions)} # 计算论文摘要TF-IDF仅用标题摘要前500字符 abstracts [p[title] p[abstract][:500] for p in papers] self.tfidf TfidfVectorizer(max_features5000, stop_wordsenglish, ngram_range(1,2)) self.paper_features self.tfidf.fit_transform(abstracts).toarray() # shape: (N_papers, 5000) def build_edge_index(self) - Dict[str, torch.Tensor]: 构建按关系类型分组的边索引 edges defaultdict(list) # CITES边paper - paper for p in self.papers: src_id self.paper2id[p[paper_id]] for ref_id in p.get(references, []): if ref_id in self.paper2id: tgt_id self.paper2id[ref_id] edges[CITES].append([src_id, tgt_id]) # AUTHORED_BY边paper - author for p in self.papers: src_id self.paper2id[p[paper_id]] for author in p.get(authors, []): if author in self.author2id: tgt_id self.author2id[author] edges[AUTHORED_BY].append([src_id, tgt_id]) # 转换为PyTorch张量 edge_index_dict {} for rel, edge_list in edges.items(): if edge_list: edge_tensor torch.tensor(edge_list, dtypetorch.long).t().contiguous() # 确保边是双向的对称关系如AUTHORED_BY需双向但CITES是单向 if rel in [AUTHORED_BY, AFFILIATED_WITH]: edge_tensor to_undirected(edge_tensor) edge_index_dict[rel] edge_tensor return edge_index_dict def build_pyg_data(self) - Data: 构建PyTorch Geometric Data对象 edge_index_dict self.build_edge_index() # 节点特征论文用TF-IDF作者/机构用零向量后续用R-GCN聚合得到 x_paper torch.tensor(self.paper_features, dtypetorch.float) x_author torch.zeros(len(self.authors), x_paper.shape[1]) x_inst torch.zeros(len(self.institutions), x_paper.shape[1]) # 拼接所有节点特征按类型顺序paper, author, institution x torch.cat([x_paper, x_author, x_inst], dim0) # 边索引字典key为关系名value为[2, num_edges]张量 edge_index {} for rel, idx in edge_index_dict.items(): edge_index[rel] idx return Data(xx, edge_indexedge_index, num_nodesx.size(0)) # 使用示例 papers loader.load_from_csv(sample_papers.csv) all_authors [a for p in papers for a in p.get(authors, [])] builder AcademicGraphBuilder(papers, all_authors, []) data builder.build_pyg_data() print(fGraph built: {data.num_nodes} nodes, {sum(v.size(1) for v in data.edge_index.values())} edges)参数说明TfidfVectorizer的max_features5000是经验阈值太少丢失区分度太多导致稀疏且训练慢ngram_range(1,2)保留“graph neural”这类关键二元组to_undirected()仅对AUTHORED_BY等对称关系调用CITES保持单向这是R-GCN能区分“谁引用谁”的前提x_author和x_inst初始化为零向量因为它们的语义信息将通过R-GCN层从连接的论文节点聚合而来无需预设特征。3.2 R-GCN模型实现三层结构每层专注一类关系R-GCN核心是为每种关系r定义独立的权重矩阵W^r节点i的更新公式为h_i^{(l1)} σ( Σ_{r∈R} Σ_{j∈N_r(i)} (1/|N_r(i)|) * W^r * h_j^{(l)} )其中N_r(i)是节点i在关系r下的邻居集合。我们实现一个三层R-GCN每层输出维度递减128→64→32最终论文节点的嵌入用于推荐# rgcn_model.py import torch import torch.nn as nn import torch.nn.functional as F from torch_geometric.nn import RGCNConv class RGCNRecommender(nn.Module): def __init__(self, num_node_types: int, num_relations: int, in_channels: int, hidden_channels: int, out_channels: int, num_bases: int 30): super().__init__() self.conv1 RGCNConv( in_channelsin_channels, out_channelshidden_channels, num_relationsnum_relations, num_basesnum_bases ) self.conv2 RGCNConv( in_channelshidden_channels, out_channelshidden_channels // 2, num_relationsnum_relations, num_basesnum_bases ) self.conv3 RGCNConv( in_channelshidden_channels // 2, out_channelsout_channels, num_relationsnum_relations, num_basesnum_bases ) self.dropout nn.Dropout(0.3) def forward(self, x: torch.Tensor, edge_index: Dict[str, torch.Tensor], edge_type: torch.Tensor) - torch.Tensor: # PyG的RGCNConv要求edge_index为[2, num_edges]edge_type为[num_edges] # 我们需将字典格式转换为拼接格式 edge_indices [] edge_types [] for i, (rel, idx) in enumerate(edge_index.items()): edge_indices.append(idx) edge_types.append(torch.full((idx.size(1),), i, dtypetorch.long)) if not edge_indices: return x full_edge_index torch.cat(edge_indices, dim1) full_edge_type torch.cat(edge_types, dim0) x self.conv1(x, full_edge_index, full_edge_type) x F.relu(x) x self.dropout(x) x self.conv2(x, full_edge_index, full_edge_type) x F.relu(x) x self.dropout(x) x self.conv3(x, full_edge_index, full_edge_type) return x # 初始化模型假设3种关系CITES, AUTHORED_BY, APPLIED_TO model RGCNRecommender( num_node_types3, # paper, author, institution num_relations3, # CITES, AUTHORED_BY, APPLIED_TO in_channels5000, # TF-IDF特征维数 hidden_channels128, out_channels32, # 最终嵌入维度 num_bases30 # 关系基数量平衡表达力与参数量 )逻辑说明num_bases30是关键超参它将每个关系r的权重矩阵W^r分解为W^r Σ_k c_k^r * B_k其中B_k是共享基矩阵c_k^r是关系特定系数。这大幅减少参数量从num_relations * in_dim * out_dim降至num_bases * in_dim * out_dim避免在小规模学术图谱上过拟合三层结构设计为“宽→窄”第一层捕获粗粒度关系如所有引用第三层聚焦细粒度语义如特定方法的应用场景符合认知渐进规律edge_type张量必须与edge_index严格对齐torch.cat拼接时顺序必须与edge_index.keys()顺序一致否则关系错位会导致训练崩溃。3.3 推荐生成基于嵌入相似度与路径可信度的双通道打分单纯用余弦相似度排序会推荐大量“同质化”论文如都讲BERT微调。我们引入路径可信度Path Confidence作为第二通道对候选论文q计算其与目标论文p之间所有长度≤3的路径如p -CITES- r -APPLIED_TO- q每条路径的权重为各边置信度乘积边置信度来自Neo4j中该关系的统计频率最终推荐分 0.7 * cos_sim 0.3 * path_confidence。# recommender.py import numpy as np from sklearn.metrics.pairwise import cosine_similarity from neo4j import GraphDatabase class HybridRecommender: def __init__(self, model, data, driver, paper2id): self.model model self.data data self.driver driver self.paper2id paper2id self.id2paper {v: k for k, v in paper2id.items()} def get_embedding_similarity(self, target_id: int, candidate_ids: List[int]) - np.ndarray: 获取目标论文与候选论文的嵌入余弦相似度 with torch.no_grad(): z self.model(self.data.x, self.data.edge_index, self.data.edge_type) target_emb z[target_id].cpu().numpy().reshape(1, -1) candidate_embs z[candidate_ids].cpu().numpy() sims cosine_similarity(target_emb, candidate_embs)[0] return sims def get_path_confidence(self, target_paper_id: str, candidate_paper_id: str) - float: 查询Neo4j计算两篇论文间最短路径的置信度 query MATCH path shortestPath((p1:Paper {paper_id: $target_id})-[*..3]-(p2:Paper {paper_id: $candidate_id})) WITH path, relationships(path) as rels UNWIND rels as r WITH path, r, CASE type(r) WHEN CITES THEN 0.95 WHEN AUTHORED_BY THEN 0.85 WHEN APPLIED_TO THEN 0.90 ELSE 0.7 END as conf RETURN path, reduce(acc 1.0, c IN collect(conf) | acc * c) as path_conf ORDER BY path_conf DESC LIMIT 1 with self.driver.session() as session: result session.run(query, target_idtarget_paper_id, candidate_idcandidate_paper_id) record result.single() return float(record[path_conf]) if record else 0.0 def recommend(self, target_paper_id: str, top_k: int 10) - List[Dict]: 混合推荐主函数 if target_paper_id not in self.paper2id: return [] target_id self.paper2id[target_paper_id] all_paper_ids list(self.paper2id.values()) # 过滤掉自身和已引用论文 candidate_ids [i for i in all_paper_ids if i ! target_id and self.data.edge_index.get(CITES, torch.empty(0)).size(1) 0 or not any(self.data.edge_index[CITES][1] i)] # 获取嵌入相似度 sim_scores self.get_embedding_similarity(target_id, candidate_ids) # 获取路径置信度对top 50候选计算避免全量查询耗时 top_candidate_ids np.array(candidate_ids)[np.argsort(sim_scores)[-50:]] path_scores [] for cid in top_candidate_ids: pid self.id2paper[cid] conf self.get_path_confidence(target_paper_id, pid) path_scores.append(conf) # 混合打分 final_scores 0.7 * sim_scores[np.argsort(sim_scores)[-50:]] 0.3 * np.array(path_scores) # 返回top_k top_indices np.argsort(final_scores)[-top_k:][::-1] recommendations [] for idx in top_indices: paper_id self.id2paper[top_candidate_ids[idx]] recommendations.append({ paper_id: paper_id, score: float(final_scores[idx]), similarity: float(sim_scores[np.argsort(sim_scores)[-50:]][idx]), path_confidence: float(path_scores[idx]) }) return recommendations # 使用示例 driver GraphDatabase.driver(bolt://localhost:7687, auth(neo4j, password)) recommender HybridRecommender(model, data, driver, builder.paper2id) recs recommender.recommend(10.1145/3394486.3403170, top_k5) for r in recs: print(fPaper {r[paper_id]}: score{r[score]:.3f} (sim{r[similarity]:.3f}, path{r[path_confidence]:.3f}))参数说明shortestPath限定长度[*..3]是性能与效果的平衡长度1直接引用太窄长度4以上路径爆炸且语义模糊关系置信度赋值CITES: 0.95基于经验统计引用关系在学术中最具权威性APPLIED_TO次之需人工标注验证AUTHORED_BY因作者合作泛滥而略低reduce(acc 1.0, c IN collect(conf) | acc * c)是Cypher的累乘函数将路径上各边置信度相乘体现“木桶效应”——任一边低置信整条路径分就低。4. 可视化分析系统PyQt6构建的交互式知识图谱探针GUI不是锦上添花而是知识图谱价值释放的关键接口。网页版可视化如Neo4j Browser适合探索但无法深度集成推荐算法与本地数据而PyQt6构建的桌面应用能实现毫秒级响应、离线运行、与推荐引擎无缝耦合且可打包为单文件exe供实验室全员使用。4.1 主窗口架构中心图谱画布 三栏控制面板我们采用QMainWindow作为主框架布局为中央QGraphicsView本文还有配套的精品资源点击获取
返回列表