1. 为什么我要把这三类基准题单独拎出来跑一遍DeepSeek-R1 的评测表里GPQA-Diamond、SimpleQA、FRAMES 这三项放在一起看很有意思它们分别代表博士级科学推理、短事实问答、跨文档多跳检索。很多开发者看到 71.5、30.1、82.5 这几个数字第一反应是“分数挺高”但真到自己复现时往往卡在“题目长什么样、怎么喂给模型、答案怎么判”这三步上。我这次的目标很明确用同一套 API 通道把三类基准各挑一道代表题跑通记录请求参数、返回内容、判定结果并给出可复制的配置骨架。适合已经拿到 Key、想自己搭评测脚本的开发者如果你还没配好通道第 2 节会先把这一步补齐。需要提前说明的是基准分数是论文/官方报告里的统计结果我们本地单题复现只能验证“通道能跑、格式能对、判定逻辑能写”不能拿单题结果去反推整体准确率。这一点想清楚后面的操作就不会跑偏。2. TaoToken 统一 Key 与 API 通道配置2.1 为什么用统一通道跑评测三类基准题的调用方式其实一样都是 chat completions 接口差别只在 prompt 构造和答案判定。如果每个模型、每个基准都换一套 Key 和 Base URL脚本会变得很难维护。TaoToken 的做法是给一个统一入口模型名通过参数切换这样同一份评测脚本改一行 model 字段就能换模型。官网入口在这里https://taotoken.net/?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content API 根地址是 https://taotoken.net/api 注意 API 地址不带 UTM 参数配置时直接写这个就行。2.2 拿 Key 与写入 settings.json先在控制台创建 API Key地址https://taotoken.net/console?utm_sourcetaotoken_aicg_blog_endutm_contentconsoleutm_campaignrewrite 。创建后复制那串 sk- 开头的字符串只显示一次建议直接写进环境变量而不是硬编码。如果你用 VS Code 系插件或兼容 OpenAI 格式的客户端settings.json 可以这样写{ ai.provider: openai-compatible, ai.baseUrl: https://taotoken.net/api, ai.apiKey: sk-你的Key, ai.model: deepseek-r1, ai.temperature: 0.6, ai.topP: 0.95, ai.maxTokens: 32768 }这里的 temperature 0.6、top-p 0.95、max_tokens 32768 是参考 DeepSeek-R1 评测时的采样设置复现基准题时保持一致结果才有可比性。2.3 config.toml 版本给 Python 脚本用如果你用 Python 直接调config.toml 更清爽[api] base_url https://taotoken.net/api api_key sk-你的Key model deepseek-r1 [sampling] temperature 0.6 top_p 0.95 max_tokens 32768 n 1注意 n1论文里 pass1 是生成 64 个响应估计的本地复现单题时先用 n1 验证通道别一上来就开 64既慢又费额度。3. 三类基准题的 prompt 构造与可复制调用3.1 GPQA-Diamond博士级物理题GPQA-Diamond 的题目特点是选项固定、答案唯一、需要多步推导。原题示例是一道自旋期望值计算正确答案是 -0.7。构造 prompt 时要把题干和选项完整给出并要求模型把最终答案放进 \boxed{}方便正则提取。import openai client openai.OpenAI( base_urlhttps://taotoken.net/api, api_keysk-你的Key ) gpqa_prompt A spin-half particle is in a linear superposition 0.5|up sqrt(3)/2|down of its spin-up and spin-down states. If |up and |down are the eigenstates of sigma_z, then what is the expectation value, up to one decimal place, of the operator 10*sigma_z 5*sigma_x? Put your final answer in \\boxed{}. resp client.chat.completions.create( modeldeepseek-r1, messages[{role: user, content: gpqa_prompt}], temperature0.6, top_p0.95, max_tokens32768 ) print(resp.choices[0].message.content)跑完你会看到模型给出推导过程最后一行是 \boxed{-0.7}。判定逻辑就是正则抓 boxed 内容和标准答案字符串比对。3.2 SimpleQA短事实问答SimpleQA 的判定规则比 GPQA 严格答案必须完全包含参考答案且不能矛盾。比如“2022 年荷兰对阿根廷比赛中哪位荷兰球员进球”标准答案是 Wout Weghorst如果模型答“Virgil van Dijk and Wout Weghorst”就算错误因为引入了矛盾信息。simpleqa_prompt Which Dutch player scored an open-play goal in the 2022 Netherlands vs Argentina game in the mens FIFA World Cup? Answer with the players name only. resp client.chat.completions.create( modeldeepseek-r1, messages[{role: user, content: simpleqa_prompt}], temperature0.6, top_p0.95, max_tokens32768 ) print(resp.choices[0].message.content)判定时建议写一个三分类函数完全包含参考答案且无矛盾 → Correct出现矛盾实体 → Incorrect没给出完整答案 → Not attempted。这样和官方分类器逻辑对齐。3.3 FRAMES跨文档多跳推理FRAMES 的题目往往需要串联多个维基页面。示例题是“未来妻子的名字”需要先查第 15 任第一夫人的母亲名字再查第二位被暗杀总统的母亲的娘家姓最后组合。正确答案是 Jane Ballou。frames_prompt If my future wife has the same first name as the 15th first lady of the United States mother and her surname is the same as the second assassinated presidents mothers maiden name, what is my future wifes name? Give the final name only. resp client.chat.completions.create( modeldeepseek-r1, messages[{role: user, content: frames_prompt}], temperature0.6, top_p0.95, max_tokens32768 ) print(resp.choices[0].message.content)这道题在官方记录里 R1 答错了本地复现时如果也答错说明通道和 prompt 都没问题属于模型能力边界不用怀疑配置。4. 验证请求与结果对照4.1 用 curl 做最小连通性验证在跑完整脚本前先用一条 curl 确认通道通curl https://taotoken.net/api/chat/completions \ -H Authorization: Bearer sk-你的Key \ -H Content-Type: application/json \ -d { model: deepseek-r1, messages: [{role: user, content: 11?}], max_tokens: 64 }返回 JSON 里有 choices[0].message.content 就说明 Key 和地址都对。如果返回 401检查 Key 是否复制完整返回 404检查 base_url 是否多写了 /v1。4.2 三类题的结果对照表基准题目类型标准答案判定方式本地复现关注点GPQA-Diamond物理推导-0.7boxed 提取推导步骤是否完整SimpleQA短事实Wout Weghorst三分类是否引入矛盾实体FRAMES多跳检索Jane Ballou精确匹配中间步骤是否可追溯跑完后把模型输出、提取答案、判定结果写进 CSV方便后续批量跑。单题验证通过后再考虑把 n 调到 64 做 pass1 估计。5. 本篇常见错排查5.1 返回内容被截断max_tokens 设太小是常见原因。GPQA 这类题推导过程长32768 是评测时的设置本地如果只给 2048模型可能还没写到 boxed 就断了。检查 finish_reason 字段如果是 length就调大 max_tokens。5.2 boxed 提取失败模型有时会写 \boxed{ -0.7 } 带空格或者用 \boxed{-0.7}。正则要写成 \boxed{\s*([^}]?)\s*}抓完再 strip。如果模型压根没写 boxed说明 prompt 里的格式要求不够强可以在末尾加一句“Final answer must be in \boxed{}”。5.3 SimpleQA 判定过严或过松官方分类器是语义级的本地用字符串包含会偏严。比如模型答“Wout Weghorst scored twice”字符串包含 Wout Weghorst应该判 Correct。但如果答“I think it was Weghorst”虽然包含官方可能判 Not attempted。建议先按包含逻辑跑再人工抽查几条校准。5.4 FRAMES 多跳题超时这类题模型内部推理链长响应时间可能到几十秒。客户端要设 timeoutPython SDK 默认 600 秒够用但如果你用 requests 手写记得加 timeout120。另外别在循环里串行跑几百条容易触发限流加个 sleep 或并发控制。6. 后续怎么把这套骨架用起来单题跑通后最实用的做法是把三类题的 prompt 模板、判定函数、结果记录抽成一个 evaluator.py模型名做成参数。这样你想对比 R1 和其他模型只改 model 字段就行。长期做编码类评测或 Agent 任务的话可以考虑 Coding Plan入口https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding-planutm_campaignrewrite 它更适合持续性的代码生成评测场景。如果只是想快速验证某个模型对某道题的回答直接用模型对话页面更省事https://taotoken.net/model-chat?utm_sourcetaotoken_aicg_blog_endutm_contentmodel-chatutm_campaignrewrite 。接入文档在 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite Key 管理在 https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi-keysutm_campaignrewrite 。我自己的习惯是新基准先跑 3 道题验证通道和判定逻辑确认无误再批量。批量时把 temperature 固定 0.6、top_p 0.95结果才和论文里的设置对齐。这套流程跑下来GPQA 和 SimpleQA 基本能稳定复现FRAMES 的多跳题偶尔会错属于正常范围记录好就行。