
简介本资源是一份面向高校计算机专业学生与Python初学者的完整课程设计项目聚焦电影数据全链路处理实践涵盖猫眼电影网站数据爬取、清洗分析及多维度可视化展示可直接用于期末大作业或课程设计并获得高分参考。压缩包共42个文件以11个核心Python脚本含爬虫、数据分析、Flask后端及可视化模块为主干辅以5个HTML前端页面、5个JS交互脚本、1个SQL数据库文件及详细README文档结构清晰、模块解耦支持开箱即用与二次开发配套9个pyc缓存文件与5个XML配置文件便于环境适配整体仅422KB轻量易部署。已有837人学习下载项目获97分高分评价代码全程中文注释附带完整运行说明与技术文档小白可快速上手进阶者亦能基于现有Flask框架与Mako模板拓展API接口或增加图表类型。1. 猫眼电影数据爬取不是“写个requests就完事”它是一套闭环的课程设计级工程含反爬绕过、结构化清洗、多维分析与ECharts动态看板适合Python入门后想交出高分作业的学生和想快速复现真实场景的数据新人你可能试过用requests.get()抓猫眼首页结果返回一堆空 div 或者 403也可能跑通了爬虫但发现票房字段全是“暂无”评分字段全是“—”时间字段全是“上映中”——这不是代码写错了是没理解猫眼的动态渲染接口分流字段脱敏三重机制。这个97分课程设计项目恰恰把这三道坎全踩实了它不用 Selenium 暴力模拟而是精准定位猫眼 Web API 的加密参数生成逻辑不靠正则硬扒 HTML而是用jsonpath解析真实响应体清洗阶段专门处理“2.3亿”→23000000、“12.5万观影人次”→125000 这类非标数字可视化不是静态图而是用 Flask ECharts 渲染带下拉筛选、时间轴联动、票房/评分/热度三指标同屏对比的交互式看板。它不是玩具 demo而是能直接放进课程设计答辩 PPT、附录里贴出可运行截图、老师扫码就能验证的完整工程。如果你正在赶期末大作业、需要一份“有细节、有注释、有文档、有避坑说明”的 Python 数据项目它比网上零散的爬虫教程更接近真实开发节奏——毕竟97分背后是三次推倒重写的反爬策略和七版迭代的字段映射表。2. 爬取层绕过猫眼动态加密参数与频率限制用 requests execjs 复现前端 sign 生成逻辑猫眼电影列表页如https://maoyan.com/films?showType1看似是静态页面实则所有关键数据影片 ID、名称、评分、票房、上映时间都由https://maoyan.com/ajax/movieOnInfoList这类接口返回且请求头中必须携带X-Requested-With: XMLHttpRequestURL 参数中必须包含uuid设备唯一标识、offset分页偏移、limit单页数量最关键的是sign字段——它是对timestamp offset limit组合字符串做 MD5 加密后再拼接特定 salt 的结果而 salt 值藏在前端 JS 文件里。项目没用 Selenium 启动浏览器而是用execjs调用本地 Node.js 环境执行猫眼官网加载的common.js片段提取 salt 并复现 sign 生成函数。这种做法比硬编码 salt 更鲁棒也比频繁换 User-Agent 更可持续。2.1 提取猫眼前端加密 salt 的真实路径与解析逻辑猫眼官网 JS 资源路径并非固定项目通过分析https://maoyan.com/首页 HTML 中script标签的src属性定位到类似/js/common.8a3b1c2d.js的文件hash 值随版本变化。下载该 JS 后用正则匹配var salt ([a-zA-Z0-9]);或const salt ([a-zA-Z0-9]);模式提取 salt 值。注意salt 不是全局常量而是随 JS 版本更新而变更因此项目在utils/crypt.py中封装了get_salt_from_js()函数自动抓取并缓存最新 salt避免手动更新。# utils/crypt.py import re import requests from bs4 import BeautifulSoup def get_salt_from_js(): 从猫眼首页动态获取最新 salt 值 headers { User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36 } # 第一步获取首页 HTML home_resp requests.get(https://maoyan.com/, headersheaders, timeout10) home_resp.raise_for_status() # 第二步解析 script 标签找到 common.js 路径 soup BeautifulSoup(home_resp.text, html.parser) script_tags soup.find_all(script, srcTrue) common_js_url None for tag in script_tags: if common in tag[src]: common_js_url https: tag[src] if tag[src].startswith(//) else tag[src] break if not common_js_url: raise ValueError(未找到 common.js 资源路径) # 第三步下载 common.js 并提取 salt js_resp requests.get(common_js_url, headersheaders, timeout10) js_resp.raise_for_status() # 正则匹配 salt 定义兼容 var/const/let salt_match re.search(r(?:var|const|let)\ssalt\s*\s*[\]([^\])[\];, js_resp.text) if not salt_match: raise ValueError(JS 文件中未找到 salt 定义) return salt_match.group(1)提示此函数需在首次运行时调用结果建议存入config.py的SALT变量中后续直接读取避免每次启动都请求猫眼首页——既降低被风控概率也提升启动速度。2.2 用 execjs 复现前端 sign 生成函数替代硬编码 MD5猫眼 sign 生成逻辑为sign md5(timestamp offset limit salt)其中 timestamp 是毫秒级时间戳精确到秒即可offset 和 limit 是整数。项目未用 Python 的hashlib.md5()手动拼接而是将前端 JS 中的 sign 函数通常形如function getSign(t, o, l) { return md5(t o l salt); }提取出来用execjs在 Python 中执行。这样做的好处是当猫眼某天改用sha256或加入随机因子时只需替换 JS 片段无需重写 Python 逻辑。# utils/crypt.py import execjs import time import hashlib # 从 cat-eye-master/static/js/sign_generator.js 中读取前端 sign 函数定义 # 该文件是项目预置的、已提取并适配的 JS 片段内容类似 # const salt abc123; # function getSign(timestamp, offset, limit) { # return md5(timestamp offset limit salt); # } # 注意md5 函数需在 JS 中自行实现或引入 crypto-js项目已内置精简版 def generate_sign(timestamp, offset, limit): 调用 JS 环境生成 sign try: # 加载预置的 sign_generator.js with open(static/js/sign_generator.js, r, encodingutf-8) as f: js_code f.read() ctx execjs.compile(js_code) return ctx.call(getSign, str(timestamp), str(offset), str(limit)) except Exception as e: # 回退到 Python 原生 MD5仅用于调试生产环境应确保 JS 存在 raw_str f{timestamp}{offset}{limit}{SALT} return hashlib.md5(raw_str.encode()).hexdigest() # 示例调用 ts int(time.time() * 1000) // 1000 # 秒级时间戳 sign generate_sign(ts, 0, 10) # 获取前10部电影 print(f生成 sign: {sign})注意execjs需要系统已安装 Node.jsWindows 下推荐使用 nvm-windows 管理多版本且sign_generator.js必须保证语法兼容性避免使用 ES6 新特性。项目static/js/目录下已提供经 Babel 转译的兼容版本直接调用即可。2.3 构建带重试与降频的请求会话应对猫眼 429 和 503猫眼对单 IP 的请求频率极为敏感连续请求超过 3 次/秒大概率触发 429 Too Many Requests长时间高频访问则返回 503 Service Unavailable。项目在api_1_0/utils.py中封装了MaoyanSession类继承requests.Session内置指数退避重试最大 3 次、随机延迟0.5~2.5 秒、请求头轮换预置 5 组 UA三大机制# api_1_0/utils.py import time import random import requests from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry class MaoyanSession(requests.Session): def __init__(self): super().__init__() # 预置 UA 池模拟不同设备/浏览器 self.ua_pool [ Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36, Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36, Mozilla/5.0 (iPhone; CPU iPhone OS 16_6 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.6 Mobile/15E148 Safari/604.1, Mozilla/5.0 (Linux; Android 13; SM-S901B) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/112.0.0.0 Mobile Safari/537.36, Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36 ] # 配置重试策略 retry_strategy Retry( total3, backoff_factor1, # 第一次重试延迟 1s第二次 2s第三次 4s status_forcelist[429, 503, 502, 504], allowed_methods[HEAD, GET, OPTIONS] ) adapter HTTPAdapter(max_retriesretry_strategy) self.mount(http://, adapter) self.mount(https://, adapter) def request(self, method, url, **kwargs): # 随机设置 UA kwargs.setdefault(headers, {})[User-Agent] random.choice(self.ua_pool) # 随机延迟0.5~2.5秒 time.sleep(random.uniform(0.5, 2.5)) return super().request(method, url, **kwargs) # 使用示例 session MaoyanSession() resp session.get(https://maoyan.com/ajax/movieOnInfoList, params{uuid: xxx, offset: 0, limit: 10, sign: yyy})关键点backoff_factor1表示重试间隔为2^retry_number * backoff_factor秒即 1s、2s、4s配合随机延迟能有效规避猫眼的频率检测模型。UA 轮换不是为了伪装而是让请求特征更接近真实用户集群行为。3. 数据清洗与存储从“暂无”“上映中”到标准数值用 pandas 实现字段级规则引擎爬取到的原始 JSON 数据充满业务语义噪声“票房”字段可能是2.3亿、12.5万、暂无“上映时间”可能是2023-04-28、2023-04-28(中国大陆)、即将上映“评分”可能是9.5、—,9.5分。若直接丢进数据库或画图图表会大量报错、排序错乱、统计失真。项目在model.py中定义了MovieDataCleaner类不依赖模糊匹配或人工字典而是为每个字段配置正则表达式 转换函数 缺失值策略的三元组规则形成可插拔的清洗流水线。3.1 定义票房字段清洗规则支持“亿/万/千”单位自动换算与缺失值归零票房字段boxOffice原始值格式混乱需统一转为整型单位元。规则引擎要求匹配(\d\.?\d*)[亿|万|千]提取数字并按单位换算1亿1e81万1e41千1e3若含“暂无”“待定”等词则置为 0若为空字符串或 None则置为 0。项目将规则写死在clean_rules.py中便于测试与复用# utils/clean_rules.py import re BOX_OFFICE_RULES { pattern: r(\d\.?\d*)([亿|万|千]), converter: lambda match: float(match.group(1)) * { 亿: 1e8, 万: 1e4, 千: 1e3 }.get(match.group(2), 1), default_value: 0, preprocess: lambda x: str(x).strip() if x else } def clean_box_office(raw_value): 清洗票房字段 processed BOX_OFFICE_RULES[preprocess](raw_value) if 暂无 in processed or 待定 in processed or 未上映 in processed: return BOX_OFFICE_RULES[default_value] # 尝试匹配带单位的数值 match re.search(BOX_OFFICE_RULES[pattern], processed) if match: return int(BOX_OFFICE_RULES[converter](match)) # 尝试匹配纯数字如 123456789 try: return int(float(processed.replace(,, ))) except (ValueError, TypeError): return BOX_OFFICE_RULES[default_value] # 测试 print(clean_box_office(2.3亿)) # 230000000 print(clean_box_office(12.5万)) # 125000 print(clean_box_office(暂无)) # 0 print(clean_box_office(123,456.78)) # 123456注意preprocess函数负责统一字符串格式去空格、转字符串converter是闭包函数避免在正则匹配外重复计算单位系数。这种解耦设计让新增字段如“观影人次”只需复制规则模板修改 pattern 和 converter 即可。3.2 上映时间字段清洗提取标准日期并标记状态类型已上映/即将上映/未定上映时间showTime需拆解为两个维度1标准datetime.date对象用于时间序列分析2状态标签released/upcoming/unknown用于后续分组统计。项目用dateutil.parser.parse()尝试解析失败则用正则提取YYYY-MM-DD再根据是否早于当前日期判断状态# utils/clean_rules.py from datetime import date, datetime from dateutil import parser import re def clean_show_time(raw_value): 清洗上映时间字段返回 (date_obj, status_tag) 元组 if not raw_value: return None, unknown raw_str str(raw_value).strip() # 规则1匹配 YYYY-MM-DD 格式 date_match re.search(r(\d{4})-(\d{1,2})-(\d{1,2}), raw_str) if date_match: try: d date(int(date_match.group(1)), int(date_match.group(2)), int(date_match.group(3))) status released if d date.today() else upcoming return d, status except (ValueError, TypeError): pass # 规则2尝试用 dateutil 解析兼容 2023年4月28日 等中文格式 try: parsed parser.parse(raw_str, fuzzyTrue) d parsed.date() status released if d date.today() else upcoming return d, status except (parser.ParserError, ValueError, TypeError): pass # 规则3关键词匹配 if any(kw in raw_str for kw in [即将上映, 预售中, 点映中]): return None, upcoming if any(kw in raw_str for kw in [未定, 待定, 暂无]): return None, unknown return None, unknown # 示例 print(clean_show_time(2023-04-28)) # (datetime.date(2023, 4, 28), released) print(clean_show_time(2023年4月28日)) # (datetime.date(2023, 4, 28), released) print(clean_show_time(即将上映)) # (None, upcoming) print(clean_show_time(2024-12-31(中国大陆))) # (datetime.date(2024, 12, 31), upcoming)关键点fuzzyTrue让dateutil.parser能容忍中文括号、空格等噪声状态标签upcoming和released直接用于后续groupby分析比用布尔值更语义清晰。3.3 评分字段清洗统一为浮点数处理“—”“暂无”及小数点格式异常评分score字段常见问题9.5→9.59.5分→9.5—→None9,5欧洲格式→9.5。项目采用两步法先用正则提取所有数字和小数点再转换为 float失败则返回 None# utils/clean_rules.py import re def clean_score(raw_value): 清洗评分字段 if not raw_value: return None raw_str str(raw_value).strip() # 移除所有非数字、非小数点字符保留 9.5、9,5 中的数字和分隔符 cleaned re.sub(r[^\d.,], , raw_str) if not cleaned: return None # 统一逗号为小数点处理欧洲格式 cleaned cleaned.replace(,, .) # 提取第一个匹配的浮点数模式如 9.5.2 只取 9.5 score_match re.search(r(\d\.\d|\d), cleaned) if score_match: try: return float(score_match.group(1)) except (ValueError, TypeError): pass return None # 测试 print(clean_score(9.5)) # 9.5 print(clean_score(9.5分)) # 9.5 print(clean_score(9,5)) # 9.5 print(clean_score(—)) # None print(clean_score(暂无)) # None print(clean_score(9.5.2)) # 9.5提示re.search(r(\d\.\d|\d)优先匹配带小数点的数再匹配整数避免9被误认为9.。返回None而非0是因为评分缺失与评分为 0 语义完全不同后续分析中None会被 pandas 自动忽略如mean()不计入而0会拉低均值。4. 数据分析层用 pandas 实现票房-评分-热度三维交叉分析识别“叫好不叫座”与“叫座不叫好”影片清洗后的数据存入 SQLitecat_eye.db表结构为movies(id, name, score, box_office, show_time, status, actors, directors, duration)。项目在script.py中构建了 5 个核心分析模块全部基于 pandas DataFrame 操作不依赖 SQL 复杂查询确保小白也能看懂每行代码的业务含义。重点不是“算出均值”而是用数据回答具体业务问题哪些影片票房高但评分低哪些影片评分高但票房差不同类型影片的平均票房分布如何4.1 识别“叫好不叫座”影片评分 ≥8.5 且票房 5000 万这是课程设计中最常被问到的问题。项目用布尔索引 query()方法实现结果 DataFrame 包含影片名、评分、票房、上映日期并按票房升序排列方便找出最典型的案例# script.py import pandas as pd import sqlite3 conn sqlite3.connect(cat_eye.db) df pd.read_sql_query(SELECT * FROM movies WHERE score IS NOT NULL AND box_office IS NOT NULL, conn) # 识别“叫好不叫座”评分≥8.5 且票房5000万50000000 good_but_low_box df.query(score 8.5 and box_office 50000000).sort_values(box_office).head(10) print(【叫好不叫座】Top 10评分≥8.5票房5000万) print(good_but_low_box[[name, score, box_office, show_time]].to_string(indexFalse)) # 输出示例 # 【叫好不叫座】Top 10评分≥8.5票房5000万 # name score box_office show_time # 《宇宙探索编辑部》 8.7 48200000 2023-04-01 # 《人生大事》 7.3 171000000 2022-06-24 # ...注意此处仅为示意实际数据需运行后查看逻辑说明query()比df[(df.score8.5) (df.box_office5e7)]更易读sort_values(box_office)让最低票房的影片排在最前突出“叫好却卖不动”的极端案例head(10)限制输出长度避免刷屏。4.2 识别“叫座不叫好”影片票房 ≥5 亿且评分 ≤6.5与上一节对称此分析关注商业成功但口碑平庸的影片。项目额外增加了“票房/评分比值”列量化“每一分评分对应的票房额”比单纯看票房更反映市场接受度# script.py # 识别“叫座不叫好”票房≥5亿500000000且评分≤6.5 big_but_low_score df.query(box_office 500000000 and score 6.5).copy() big_but_low_score[box_per_score] big_but_low_score[box_office] / big_but_low_score[score] big_but_low_score big_but_low_score.sort_values(box_per_score, ascendingFalse).head(10) print(\n【叫座不叫好】Top 10票房≥5亿评分≤6.5按票房/评分比值排序) print(big_but_low_score[[name, score, box_office, box_per_score, show_time]].to_string( indexFalse, formatters{box_office: {:,.0f}.format, box_per_score: {:,.0f}.format} )) # 输出示例格式化后 # name score box_office box_per_score show_time # 《满江红》 7.0 4,540,000,000 648,571,429 2023-01-22 # 《流浪地球2》 7.7 4,020,000,000 522,077,922 2023-01-22 # ...注意实际数据中需满足 score6.5 条件参数说明formatters参数让大数字自动添加千位分隔符4,540,000,000提升可读性box_per_score是核心指标值越大说明“观众用真金白银投票但专业评分不高”典型如贺岁档合家欢电影。4.3 类型片票房分布分析用 value_counts histplot 揭示“喜剧片是票房基本盘”猫眼数据中无直接“类型”字段但项目从actors和directors字段的文本中提取关键词如“开心麻花”“宁浩”指向喜剧“郭帆”“吴京”指向科幻/动作构建简易类型标签。然后用groupby().agg()计算各类型平均票房、中位数票房、影片数量# script.py # 简易类型标注基于导演/主演关键词 def guess_genre(row): director str(row.get(directors, )).lower() actor str(row.get(actors, )).lower() if 开心麻花 in director or 宁浩 in director or 黄渤 in actor or 沈腾 in actor: return comedy elif 郭帆 in director or 吴京 in actor or 科幻 in str(row.get(name, )): return sci_fi elif 张艺谋 in director or 陈凯歌 in director or 文艺 in str(row.get(name, )): return art_film else: return other df[genre] df.apply(guess_genre, axis1) # 各类型票房统计 genre_stats df.groupby(genre).agg({ box_office: [count, mean, median], score: [mean, std] }).round(2) print(\n【类型片票房与评分统计】单位元) print(genre_stats) # 输出示例 # box_office score # count mean median mean std # genre # art_film 12 8.2e06 5.1e06 7.2 0.82 # comedy 45 1.2e08 9.5e07 6.8 0.91 # sci_fi 18 3.4e08 2.8e08 7.1 0.75 # other 120 2.1e07 1.3e07 6.5 1.02关键点agg()一次性计算多个统计量避免多次groupbyround(2)让浮点数显示简洁count列暴露数据量差异喜剧片数量最多解释为何其平均票房虽非最高却是“基本盘”。5. 数据可视化层FlaskECharts 构建可交互的电影数据看板支持下拉筛选与时间轴联动项目最终交付物不是几张静态 PNG而是一个本地运行的 Web 看板http://127.0.0.1:5000。它用 Flask 提供 API 接口/api/movies返回 JSON前端用 ECharts 渲染折线图票房趋势、柱状图类型分布、散点图评分 vs 票房、饼图状态占比所有图表支持鼠标悬停查看详情、点击图例开关系列、拖拽时间轴缩放。这不是炫技而是课程设计答辩时老师最想看到的“可演示、可交互、可验证”成果。5.1 Flask 后端提供标准化 JSON 接口支持分页与条件过滤manager.py是 Flask 入口定义了/api/movies接口接收page、per_page、statusreleased/upcoming、min_score等参数返回符合筛选条件的影片列表。关键点在于不返回原始数据库字段而是返回清洗后、可直接用于 ECharts 的结构化 JSON# manager.py from flask import Flask, request, jsonify import pandas as pd import sqlite3 app Flask(__name__) app.route(/api/movies) def get_movies(): page int(request.args.get(page, 1)) per_page int(request.args.get(per_page, 20)) status request.args.get(status) # released, upcoming, all min_score float(request.args.get(min_score, 0)) conn sqlite3.connect(cat_eye.db) # 构建动态 SQL 查询 query SELECT name, score, box_office, show_time, status, genre FROM movies WHERE 11 params [] if status and status ! all: query AND status ? params.append(status) if min_score 0: query AND score ? params.append(min_score) # 分页 offset (page - 1) * per_page query LIMIT ? OFFSET ? params.extend([per_page, offset]) df pd.read_sql_query(query, conn, paramsparams) conn.close() # 转为 ECharts 兼容格式数值字段保持原样日期转字符串 result df.to_dict(records) for item in result: if item[show_time]: item[show_time] item[show_time].strftime(%Y-%m-%d) if hasattr(item[show_time], strftime) else str(item[show_time]) return jsonify({ data: result, total: len(df), page: page, per_page: per_page }) if __name__ __main__: app.run(debugTrue)注意strftime(%Y-%m-%d)确保日期字段是字符串而非datetime.date对象避免 ECharts 解析失败to_dict(records)生成标准 JSON 数组无需额外序列化。5.2 ECharts 前端用 option 配置实现票房趋势折线图与评分-票房散点图联动templates/index.html中嵌入两个 ECharts 实例chart1票房趋势监听chart2散点图的点击事件反之亦然。联动逻辑在static/js/main.js中实现核心是chart.on(click, ...)和chart.dispatchAction({ type: highlight, ... })// static/js/main.js // 初始化票房趋势折线图 const chart1 echarts.init(document.getElementById(trend-chart)); chart1.setOption({ title: { text: 近30日热门影片票房趋势 }, tooltip: { trigger: axis }, xAxis: { type: category, data: [] }, // x轴为影片名 yAxis: { type: value, name: 票房万元 }, series: [{ name: 票房, type: line, data: [], smooth: true, emphasis: { focus: series } }], legend: { data: [票房] } }); // 初始化评分-票房散点图 const chart2 echarts.init(document.getElementById(scatter-chart)); chart2.setOption({ title: { text: 评分 vs 票房散点图 }, tooltip: { trigger: item, formatter: {a}br/{b}: {c[0]}分, {c[1]}万元 }, xAxis: { type: value, name: 评分 }, yAxis: { type: value, name: 票房万元 }, series: [{ name: 影片, type: scatter, data: [], // 格式[[score, box_office/10000, name], ...] symbolSize: function (data) { return Math.sqrt(data[1]) / 5; }, // 票房越大点越大 emphasis: { focus: self } }] }); // 散点图点击时高亮趋势图中对应影片 chart2.on(click, function (params) { const clickedName params.data[2]; chart1.dispatchAction({ type: highlight, seriesIndex: 0, dataIndex: chart1.getDataSet().findIndex(d d.name clickedName) }); }); // 趋势图点击时高亮散点图中对应影片 chart1.on(click, function (params) { const clickedName params.name; const scatterData chart2.getOption().series[0].data; const dataIndex scatterData.findIndex(d d[2] clickedName); if (dataIndex ! -1) { chart2.dispatchAction({ type: highlight, seriesIndex: 0, dataIndex: dataIndex }); } });关键点symbolSize函数让散点大小与票房平方根成正比视觉上更符合认知dispatchAction({ type: highlight })是 ECharts 官方推荐的联动方式比手动 setOption 更高效formatter中的{c[0]}分, {c[1]}万元直本文还有配套的精品资源点击获取