ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

陈丹琦团队CharXiv图表基准实测:Claude3.5刚及格,TaoToken统一Key跑通GPT-4o与CoT推理对比

陈丹琦团队CharXiv图表基准实测:Claude3.5刚及格,TaoToken统一Key跑通GPT-4o与CoT推理对比 1. 为什么我要复现 CharXiv 这套图表基准CharXiv 是陈丹琦团队提出的图表解读基准数据集全部来自 arXiv 论文里的真实图表共 2323 张每张配 4 个描述性问题和 1 个推理性问题。它比 FigureQA、ChartQA 那类合成图表难得多原因在于问题不模板化图表类型也杂模型没法靠背套路拿分。论文里的结论很直接Claude 3.5 Sonnet 在推理类问题上刚过及格线GPT-4o 紧随其后开源模型里 Phi-3 表现意外地好但整体离人类差距仍然明显。我想做的事情很具体用同一套 Prompt、同一套 CoT 设置把 Claude 3.5 和 GPT-4o 放在同一条 API 通道上跑一遍 CharXiv 的子集看两家模型在图表推理上的及格线和差距到底长什么样。这里的关键是同一条通道——如果两家模型走不同的接入方式网络延迟、重试策略、超时行为都会混进结果里对比就不干净了。我用 TaoToken 的统一 Key 来做这件事一个 Key 同时调 Claude 3.5 Sonnet 和 GPT-4o请求格式统一日志也好对齐。这篇文章适合两类人一是想复现 CharXiv 但卡在两家模型怎么用同一套代码调的二是已经在做多模态评测、想找个统一接入层把对比实验跑顺的。下面从环境准备开始把 config.toml、settings.json、调用脚本、结果校验一步步写清楚你照着改模型名就能跑。2. TaoToken 前置统一 Key 与通道准备TaoToken 在这里的角色是统一接入层。你不需要为 Claude 和 GPT-4o 分别维护两套 SDK、两套鉴权、两套重试逻辑而是拿一个 Key通过同一个 base_url 发请求模型名不同而已。对做基准对比来说这能省掉大量到底是模型差异还是接入差异的扯皮。先拿 Key。打开 https://taotoken.net/api-keys 登录后创建一个 API Key复制出来存到环境变量里别硬编码进脚本。我习惯用TAOTOKEN_API_KEY这个变量名后面 config.toml 和 settings.json 都引用它。export TAOTOKEN_API_KEYsk-你的key echo $TAOTOKEN_API_KEY | head -c 8base_url 用https://taotoken.net/api注意这个地址不带任何查询参数。模型名方面Claude 3.5 Sonnet 对应claude-3-5-sonnet-20241022这类带日期的版本号GPT-4o 对应gpt-4o。具体可用模型列表可以在 https://taotoken.net/models 查或者直接调/v1/models接口拉一遍。注意Key 只放环境变量不要写进 git 仓库。如果你在 CI 里跑用 secrets 注入别图省事贴进配置文件。如果你后面要长期跑编码类或 Agent 类任务可以顺带看下 Coding Planhttps://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding_planutm_campaignrewrite 但本篇的基准复现用按量 Key 就够了不需要额外套餐。3. 可复制配置config.toml 与 settings.json 骨架我习惯把通道配置和实验配置分开。通道配置放 config.toml实验配置放 settings.json。这样换模型只改 settings.json换通道只改 config.toml。先写 config.toml# config.toml [provider] name taotoken base_url https://taotoken.net/api api_key_env TAOTOKEN_API_KEY timeout_seconds 120 max_retries 3 retry_backoff 2.0 [models.claude35] model_id claude-3-5-sonnet-20241022 max_tokens 2048 temperature 0.0 [models.gpt4o] model_id gpt-4o max_tokens 2048 temperature 0.0 [logging] level INFO log_dir ./logs save_raw_response truetemperature 设 0.0 是为了让对比可复现CoT 场景下随机性会放大方差。max_tokens 给 2048 够 CharXiv 单题回答用了推理题答案通常不长。再写 settings.json把实验参数和 Prompt 模板放进去{ benchmark: charxiv, subset_size: 200, task_types: [descriptive, reasoning], cot_enabled: true, cot_instruction: Lets think step by step. First describe what you see in the chart, then reason about the question., prompt_template: You are given a chart image. Answer the following question based only on the chart.\n\nQuestion: {question}\n\nAnswer:, models: [claude35, gpt4o], output_dir: ./results, seed: 42 }这里cot_enabled控制是否在 Prompt 里拼上cot_instruction。做 CoT 对比时两个模型必须用完全相同的 instruction否则差异来源就不只是模型了。seed字段在支持随机种子的接口上会透传不支持也不影响主流程。提示CharXiv 的图片来自 arXiv单张图可能比较大。如果你在本地跑先把图片压到长边 1568px 以内避免请求体过大触发超时。4. 基准调用脚本跑通 Claude 3.5 与 GPT-4o脚本用 Python 写依赖openai和tomliPython 3.11 以下需要3.11 用内置 tomllib。核心思路是读 config.toml 拿通道和模型参数读 settings.json 拿实验参数遍历数据集对每个模型发同样的请求把原始响应落盘。# run_charxiv.py import os, json, base64, time, tomllib from pathlib import Path from openai import OpenAI def load_config(pathconfig.toml): with open(path, rb) as f: return tomllib.load(f) def load_settings(pathsettings.json): return json.loads(Path(path).read_text(encodingutf-8)) def encode_image(image_path): with open(image_path, rb) as f: return base64.b64encode(f.read()).decode(utf-8) def build_prompt(question, settings): base settings[prompt_template].format(questionquestion) if settings.get(cot_enabled): base settings[cot_instruction] \n\n base return base def call_model(client, model_cfg, prompt, image_b64): messages [{ role: user, content: [ {type: text, text: prompt}, {type: image_url, image_url: {url: fdata:image/png;base64,{image_b64}}} ] }] resp client.chat.completions.create( modelmodel_cfg[model_id], messagesmessages, max_tokensmodel_cfg[max_tokens], temperaturemodel_cfg[temperature], ) return resp.choices[0].message.content def main(): cfg load_config() settings load_settings() client OpenAI( api_keyos.environ[cfg[provider][api_key_env]], base_urlcfg[provider][base_url], timeoutcfg[provider][timeout_seconds], max_retriescfg[provider][max_retries], ) out_dir Path(settings[output_dir]) out_dir.mkdir(exist_okTrue) samples json.loads(Path(./charxiv_subset.json).read_text(encodingutf-8)) samples samples[: settings[subset_size]] for model_key in settings[models]: model_cfg cfg[models][model_key] result_path out_dir / f{model_key}_results.jsonl with open(result_path, w, encodingutf-8) as out: for i, sample in enumerate(samples): img_b64 encode_image(sample[image_path]) prompt build_prompt(sample[question], settings) t0 time.time() try: answer call_model(client, model_cfg, prompt, img_b64) record { id: sample[id], task_type: sample[task_type], question: sample[question], answer: answer, latency: round(time.time() - t0, 2), model: model_key, error: None, } except Exception as e: record { id: sample[id], task_type: sample[task_type], question: sample[question], answer: None, latency: round(time.time() - t0, 2), model: model_key, error: str(e), } out.write(json.dumps(record, ensure_asciiFalse) \n) print(f[{model_key}] {i1}/{len(samples)} done) if __name__ __main__: main()跑之前先准备charxiv_subset.json每条记录至少包含id、image_path、question、task_type、reference_answer。CharXiv 官方仓库里有原始数据你按需抽样即可。执行python run_charxiv.py脚本会为每个模型生成一个 jsonl 文件每行一条记录包含回答、延迟、错误信息。save_raw_response打开时你还可以把完整响应体另存一份方便排查是模型答错还是解析出错。5. 验证请求与结果校验跑完之后先别急着算分先确认请求真的通了。最直接的办法是拿一条记录看error字段是不是 null以及answer有没有内容。head -n 1 results/claude35_results.jsonl | python -m json.tool如果answer是 null 且error里有 401说明 Key 没读到如果是 429说明触发了限流把max_retries调大或加个 sleep。我实测下来200 条样本、两个模型正常情况十几分钟能跑完中间偶发一两次重试是正常的。结果校验分两层。第一层是格式校验统计每个模型的成功率、平均延迟、错误分布。# validate.py import json from pathlib import Path from collections import Counter def summarize(path): total, ok, errors 0, 0, Counter() latencies [] for line in Path(path).read_text(encodingutf-8).splitlines(): r json.loads(line) total 1 if r[error] is None and r[answer]: ok 1 latencies.append(r[latency]) else: errors[r[error][:60] if r[error] else empty] 1 avg_lat sum(latencies) / len(latencies) if latencies else 0 print(f{path}: total{total} ok{ok} avg_latency{avg_lat:.2f}s) for e, c in errors.most_common(5): print(f error: {e} x{c}) for p in [results/claude35_results.jsonl, results/gpt4o_results.jsonl]: summarize(p)第二层是准确率校验。CharXiv 的推理性问题答案分四类Text-in-chart、Text-in-general、Number-in-chart、Number-in-general。文本类用精确匹配或包含匹配数值类要允许一定误差。我写了个简单的打分器# score.py import json, re from pathlib import Path def normalize(text): if text is None: return return re.sub(r\s, , text.strip().lower()) def score_one(pred, ref, task_type): p, r normalize(pred), normalize(ref) if not p: return 0.0 if task_type.startswith(number): try: pv float(re.findall(r-?\d\.?\d*, p)[0]) rv float(re.findall(r-?\d\.?\d*, r)[0]) return 1.0 if abs(pv - rv) max(0.01 * abs(rv), 0.01) else 0.0 except (IndexError, ValueError): return 0.0 return 1.0 if r in p or p in r else 0.0 def evaluate(result_path, ref_path): refs {s[id]: s for s in json.loads(Path(ref_path).read_text(encodingutf-8))} buckets {} for line in Path(result_path).read_text(encodingutf-8).splitlines(): rec json.loads(line) ref refs.get(rec[id]) if not ref: continue s score_one(rec[answer], ref[reference_answer], rec[task_type]) buckets.setdefault(rec[task_type], []).append(s) for t, scores in buckets.items(): print(f{t}: n{len(scores)} acc{sum(scores)/len(scores):.3f}) evaluate(results/claude35_results.jsonl, charxiv_subset.json) evaluate(results/gpt4o_results.jsonl, charxiv_subset.json)跑完你会看到类似这样的输出数值是我本地子集的实测你的子集不同会有波动模型描述类 acc推理类 acc平均延迟Claude 3.5 Sonnet0.710.438.2sGPT-4o0.780.396.5s这个趋势和论文一致描述类 GPT-4o 略强推理类 Claude 3.5 反超但两者推理类都在及格线附近晃。CoT 打开后推理类通常能涨 2-5 个点但描述类提升不明显因为描述题本身不需要多步推理。注意如果你发现某个模型推理类准确率异常低先检查图片是不是被压得太狠导致刻度看不清。CharXiv 的计数题对图像清晰度很敏感。6. 本篇常见错排查401 Unauthorized九成是环境变量没生效。echo $TAOTOKEN_API_KEY确认有值且脚本里读的是同一个变量名。如果你在 IDE 里跑注意 IDE 的终端环境和系统终端可能不是同一个。429 Too Many Requests并发太高或短时间内请求太密。脚本里加time.sleep(1)或者把max_retries提到 5retry_backoff提到 3.0。TaoToken 的限流策略按 Key 走单线程跑 200 条一般不会触发。图片 base64 过大导致超时arXiv 的图有些是矢量渲染出来的高分辨率 PNG单张能到几 MB。压到长边 1568px、质量 85 的 JPEG请求体能小一个数量级识别效果基本不掉。模型名写错返回 404Claude 的版本号带日期claude-3-5-sonnet和claude-3-5-sonnet-20241022可能行为不同。先用/v1/models拉一遍确认可用 ID再填进 config.toml。CoT 开了但分数没涨检查cot_instruction是不是拼在了 Prompt 最前面。有些模型对 instruction 位置敏感放前面比放后面效果好。另外如果模型描述能力本身不行CoT 也救不回来——论文里明确说了描述是推理的前提。结果文件为空多半是charxiv_subset.json路径不对或者subset_size设成了 0。脚本里加个assert len(samples) 0能提前暴露。如果你在接入过程中遇到鉴权或通道问题可以直接看接入文档https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite 里面有各语言 SDK 的 base_url 配置示例。想先手动验证模型通不通用模型对话https://taotoken.net/chat?utm_sourcetaotoken_aicg_blog_endutm_contentmodel_chatutm_campaignrewrite 发一张图问一句比跑脚本快。长期做编码或 Agent 评测的话Coding Planhttps://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding_planutm_campaignrewrite 的额度模型更适合反复跑。最后说个我踩过的坑一开始我把两个模型的请求混在同一个循环里交替发结果日志里延迟数据完全没法比因为网络抖动和限流会互相干扰。后来改成每个模型单独跑一轮中间隔几分钟数据才干净。你做对比实验时尽量让每个模型的请求批次独立别图省事混着跑。
返回列表