尧图网络科技YAOTU DIGITAL 获取报价
获取报价
首页 / 资讯中心 / 文章详情

OpenCompass 评估 General365:级联数学验证与 LLM 判题完整实战指南

发布时间:2026/9/29 3:27:07

资讯中心
01
ARTICLE

OpenCompass 评估 General365:级联数学验证与 LLM 判题完整实战指南

OpenCompass 评估 General365:级联数学验证与 LLM 判题完整实战指南
模型评测人工智能大模型AI 评测【免费下载链接】opencompassOpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, etc, across 100 datasets covering knowledge, reasoning, coding, science, language, long-context, and safety.项目地址https://gitcode.com/gh_mirrors/op/opencompass点击查看免费下载General365 是由美团长光meituan-longcat发布的高难度通用问答评测集用于考察大模型在数学、文本推理与选择题等场景下的真实能力。本文以 OpenCompass 仓库中 General365 配置目录 的官方文档为核心结合 数据加载实现、级联评估器源码 与 测试用例完整讲解如何在 OpenCompass 中一键复现 General365 官方评测数学题走「规则验证 LLM 判题」级联链路文本题直接走 LLM 判题链路并复现论文中基于 GPT-4.1 判题官的结果。一、General365 评测配置概述1.1 数据集来源与配置骨架General365 配置直接通过 HuggingFace 的datasets.load_dataset加载官方公开数据集meituan-longcat/General365_Public无需本地预处理入口配置文件为 general365_rawprompt_cascade_llmjudge_gen.py。配置文件的核心结构DATASET_PATH meituan-longcat/General365_Public reader_cfg dict( input_columns[question, answer_type, float_round, id], output_columnanswer, ) infer_cfg dict( prompt_templatedict( typeRawPromptTemplate, messages[dict(roleuser, content{question})], ), retrieverdict(typeZeroRetriever), inferencerdict(typeGenInferencer), )配置将原始问题question原样送入模型生成答案采用RawPromptTemplate保留官方原始 prompt、ZeroRetriever不做示例检索、GenInferencer执行生成式推理。1.2 两条评测子任务该配置把数据集按answer_type划分为两个子任务分别注册为General365-mathverify与General365-text子任务覆盖答案类型评测方式General365-mathverifynumber数值、interval区间先MATHVerifyEvaluator规则校验失败样本再交GenericLLMEvaluator判题级联General365-texttext文本、single_choice单选、multiple_choice多选直接由GenericLLMEvaluator判题对应配置如下math_dataset_cfg dict( typeGeneral365Dataset, pathDATASET_PATH, answer_types[number, interval], reader_cfgreader_cfg, ) text_dataset_cfg dict( typeGeneral365Dataset, pathDATASET_PATH, answer_types[text, single_choice, multiple_choice], reader_cfgreader_cfg, )两份数据集配置共用同一份reader_cfg仅通过answer_types参数从原始数据中筛选各自负责的样本。二、数据加载与答案类型过滤的实现原理数据集加载由 General365Dataset 实现它通过LOAD_DATASET.register_module()注册为 OpenCompass 的可用数据集模块LOAD_DATASET.register_module() class General365Dataset(BaseDataset): staticmethod def load(path: str meituan-longcat/General365_Public, split: str test, answer_types: Optional[Iterable[str]] None, **kwargs) - Dataset: dataset load_dataset(pathpath, splitsplit, **kwargs) if answer_types is not None: answer_types set(answer_types) dataset dataset.filter( lambda item: item[answer_type] in answer_types) return dataset关键行为默认加载testsplit即meituan-longcat/General365_Public的test划分answer_types传入后会按样本的answer_type字段做集合过滤只保留需要的答案类型测试用例 test_general365_loads_huggingface_and_filters_answer_types 验证了「先调用load_dataset(path, splittest)再按answer_type过滤」的完整加载路径。三、级联评测规则优先、LLM 兜底的完整链路General365-mathverify是本文配置中最核心的设计采用 CascadeEvaluator 实现「先规则、后 LLM」的级联打分。3.1 配置写法eval_cfgdict( evaluatordict( typeCascadeEvaluator, rule_evaluatordict(typeMATHVerifyEvaluator), llm_evaluator_llm_evaluator(math_dataset_cfg), parallelFalse, )),其中parallelFalse表示严格级联模式只有被规则评估器判为错误correctFalse的样本才会送进 LLM 判题从而大幅节省 LLM 判题调用量。3.2 级联评估流程源码视角从 cascade_evaluator.py 的 score 方法 可以看到完整链路规则初评对每个样本先做预测后处理调用MATHVerifyEvaluator的score计算规则得分并将结果标记为evaluation_methodrule失败样本收集级联模式下仅将规则判错的样本加入failed_indices列表LLM 复评用test_set.select(failed_indices)抽取失败子集为子集补上prediction、reference列交由GenericLLMEvaluator判题结果合并最终准确率 规则正确样本数 被 LLM 改判为正确的样本数并输出cascade_stats含rule_accuracy、llm_accuracy、final_accuracy等统计每个样本同时保留rule_evaluation与llm_evaluation明细。值得一提的是CascadeEvaluator 会在${out_dir}_llm_judge路径缓存 LLM 判题结果若缓存样本数与当前需求一致则直接复用避免重复付费调用见 缓存加载逻辑。3.3 规则评估器 MATHVerifyEvaluatorMATHVerifyEvaluator 依赖math_verify与latex2sympy2_extended两个包对数值、区间答案做表达式解析与等价性验证。若环境缺少依赖会抛出如下提示pip install math_verify latex2sympy2_extended对于超时timed_out或解析出错error的样本规则评估器直接判为不正确交由 LLM 兜底。四、LLM 判题官官方 Prompt 与 GPT-4.1 复现General365-text以及级联链路中的失败样本都由GenericLLMEvaluator完成判题。判题 Prompt 完全沿用官方项目配置文件内置了系统提示词与用户提示词JUDGE_SYSTEM_PROMPT You are an expert evaluator for question answering systems. Your task is to determine if a prediction correctly answers a question based on the ground truth. Rules: 1. The prediction is correct if it captures all the key information from the ground truth. 2. The prediction is correct even if phrased differently as long as the meaning is the same. 3. The prediction is incorrect if it contains incorrect information or is missing essential details. 4. Do not challenge the correctness of the standard answer. 5. If the standard answer includes multiple possibilities, the prediction must include all and only those possibilities to be considered correct. Output a JSON object with a single field accuracy whose value is true or false. JUDGE_PROMPT Question: {question} Ground truth: {answer} Prediction: {prediction}判题官被要求输出形如{accuracy: true}的 JSON随后由 parse_general365_judgement 解析支持去除json ... 代码块包裹后解析 JSON仅当accuracy字段为布尔值时返回该值解析失败返回None统计为 parse error。4.1 配置判题官模型_llm_evaluator工厂函数统一构造 LLM 判题配置def _llm_evaluator(dataset_cfg): return dict( typeGenericLLMEvaluator, prompt_templatedict( typeRawPromptTemplate, messages[ dict(rolesystem, contentJUDGE_SYSTEM_PROMPT), dict(roleuser, contentJUDGE_PROMPT), ], ), dataset_cfgdataset_cfg, judge_cfgdict(), dict_postprocessordict(typegeneral365_llmjudge_postprocess), )judge_cfgdict()为空时GenericLLMEvaluator会回退到默认判题配置见 default_judge_cfg该配置从三个环境变量读取判题模型信息环境变量作用默认值OC_JUDGE_MODEL判题模型名称必填无OC_JUDGE_API_KEY判题 API Key必填无OC_JUDGE_API_BASE判题 API 地址可选https://api.openai.com/v1/4.2 复现论文结论的判题设置官方文档明确指出LLM 判题 Prompt 完全遵循官方项目要复现论文结果请将判题官配置为 GPT-4.1。当前官方仓库默认使用 GPT-4.1-mini而论文报告的是 GPT-4.1。因此复现论文时建议按如下方式设置环境变量export OC_JUDGE_MODELgpt-4.1 export OC_JUDGE_API_KEYyour-api-key # 可选如使用代理或自定义网关 # export OC_JUDGE_API_BASEhttps://api.openai.com/v1/GenericLLMEvaluator的默认判题配置中还包含temperature0.001、query_per_second16、batch_size1024、max_out_len16384、max_seq_len49152等参数以保证判题输出稳定且吞吐足够。五、判题结果后处理与 OpenCompass 指标对接判题官输出的是官方格式的{accuracy: bool}布尔判定需要通过 general365_llmjudge_postprocess 转换为 OpenCompass 的评分指标。该后处理器注册为DICT_POSTPROCESSORS模块逐条解析每条样本的prediction字段并产出{ accuracy: correct / total * 100 if total else 0.0, # 百分制准确率 correct_count: correct, incorrect_count: total - correct - parse_errors, parse_error_count: parse_errors, total: total, details: output, }其中每条样本还会额外写入correct布尔与judge_parsed解析出的判定两个字段方便在结果文件中排查判题失败样本。测试用例 test_general365_official_judge_postprocessor 验证了「正确样本 1、错误样本不计、解析失败计入 parse_error」的完整统计逻辑。六、汇总结果按样本数加权的官方微平均两个子任务的结果通过 General365 汇总组配置 合并为总榜分数general365_summary_groups [ dict( nameGeneral365, subsets[General365-mathverify, General365-text], # General365_Public 包含 484 个数学样本和 236 个文本样本。 # 按样本数加权可复现官方在全部 720 个样本上的微平均 # 而非简单地平均两个子集的准确率。 weights{ General365-mathverify: 484, General365-text: 236, }, ) ]要点General365_Public共720 个样本数学类 484 个、文本类 236 个官方分数为按样本数加权的微平均即(math_accuracy × 484 text_accuracy × 236) / 720而不是两个子集的简单算术平均测试用例 test_general365_summary_group_uses_sample_count_weights 以 math75.0、text50.0 为例验证了加权平均公式(75.0 * 484 50.0 * 236) / 720的计算结果。七、端到端运行与验证7.1 运行评测在安装好 OpenCompass 依赖的前提下可通过如下方式运行 General365 评测示例为单机运行集群环境请按项目 runner 文档 适配# 先配置判题官环境变量 export OC_JUDGE_MODELgpt-4.1 export OC_JUDGE_API_KEYyour-api-key # 运行评测实际以 run.py 参数为准 python run.py --config opencompass/configs/datasets/General365/general365_rawprompt_cascade_llmjudge_gen.py7.2 配置自检仓库在 test_general365_config_uses_rawprompt_and_requested_evaluators 中对配置做了完整断言可用来核对你的改动是否偏离了官方链路infer_cfg.prompt_template.type必须是RawPromptTemplatemath 子任务的 evaluator 必须是CascadeEvaluator且rule_evaluator为MATHVerifyEvaluator、llm_evaluator为GenericLLMEvaluator、parallelFalsetext 子任务的 evaluator 直接为GenericLLMEvaluatoranswer_types分别为[number, interval]与[text, single_choice, multiple_choice]。八、总结一条链路看懂 General365环节实现关键参数/文件数据加载General365Dataset.load按answer_type过滤general365.py数学题评测规则 LLM 级联parallelFalsecascade_evaluator.py文本/选择题评测直接 LLM 判题generic_llm_evaluator.py判题官官方 Prompt输出{accuracy: bool}环境变量OC_JUDGE_MODEL/OC_JUDGE_API_KEY结果解析general365_llmjudge_postprocessgeneral365.py总分汇总484/236 样本数加权微平均General365.py通过上述配置你可以在 OpenCompass 中完整复现 General365 官方的「数学规则验证 LLM 判题」评测流程若希望对齐论文数据请务必将判题官设置为 GPT-4.1而非官方仓库默认的 GPT-4.1-mini并通过OC_JUDGE_MODEL、OC_JUDGE_API_KEY环境变量完成配置。赞分享模型评测人工智能大模型AI 评测【免费下载链接】opencompassOpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, etc, across 100 datasets covering knowledge, reasoning, coding, science, language, long-context, and safety.项目地址https://gitcode.com/gh_mirrors/op/opencompass点击查看免费下载相关推荐为什么选择 Wagmi Core面向以太坊应用的响应式基础设施及其设计哲学解析为什么选择 Wagmi Core面向以太坊应用的响应式基础设施及其设计哲学解析 Wagmi Core 是以太坊应用开发的框架无关framework agno模型评测人工智能大模型AI 评测OpenCompass LLM 评判器GenericLLMEvaluator / CascadeEvaluator使用指南以 LLM 为评判器的模型评估实战OpenCompass LLM 评判器GenericLLMEvaluator / CascadeEvaluator使用指南以 LLM 为评判器的模型评估实模型评测人工智能大模型AI 评测Pydantic Evals 评估器Evaluator完全指南从确定性断言到 LLM 裁判与实验级报告评估Pydantic Evals 评估器Evaluator完全指南从确定性断言到 LLM 裁判与实验级报告评估 Pydantic Evals 是 Pydant人工智能大模型AI Agent工具调用MCP Clients上一篇pypdf 与其他 Python PDF 库的全面对比纯 Python 定位、生态脉络与选型指南下一篇CANN Runtime 异步任务并发模型全解析主机/设备并行、多 Kernel 调度与数据传输重叠创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
02
RELATED NEWS

相关资讯

更多网站建设与数字化升级内容

03
WHY YAOTU

想打造同款高转化官网?

懂行业、懂生意,从建站到增长一站式陪跑

◈

场景化定制

不做模板站,围绕你的业务场景量身设计,小众不撞款。

◐

营销型架构

以转化目标组织内容与路径,让官网真正带来询盘。

▲

全周期服务

设计、开发、运营、运维一体,上线只是开始。

免费获取你的建站方案

留下需求,专属顾问 24 小时内为你输出方案建议。