尧图网络科技YAOTU DIGITAL 获取报价
获取报价
首页 / 资讯中心 / 文章详情

OpenCompass 实战:HumanEval Pro 代码生成评测配置与结果复现指南

发布时间:2026/9/29 6:16:51

资讯中心
01
ARTICLE

OpenCompass 实战:HumanEval Pro 代码生成评测配置与结果复现指南

OpenCompass 实战:HumanEval Pro 代码生成评测配置与结果复现指南
模型评测人工智能大模型AI 评测【免费下载链接】opencompassOpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, etc, across 100 datasets covering knowledge, reasoning, coding, science, language, long-context, and safety.项目地址https://gitcode.com/gh_mirrors/op/opencompass点击查看免费下载本指南聚焦 OpenCompass 中对HumanEval Pro数据集的评测实践以仓库内置的评测结果与配置文件为骨架讲解该数据集的构造方式、评测配置的三段式写法reader / infer / eval、远程代码评估服务的调用机制以及如何在当前仓库中直接复现文中给出的 pass1 结果。读完本文你将掌握在 OpenCompass 中接入多文件代码生成数据集并完成自动评测的完整路径。HumanEval Pro 是什么HumanEval Pro 是 OpenCompass 内置的代码生成评测数据集之一其核心特点在于题目之间的依赖关系第二个问题的解法需要调用第一个问题的解法一次或多次。这与传统的单文件、单函数代码补全基准如 HumanEval不同考察的是模型在多文件协作场景下的编程能力。从仓库源码 opencompass/datasets/humaneval_pro.py 可以看到每条样本由raw_problem原始问题与new_problem新问题两部分构成评测时模型需要把两个问题的解法写在一个 Python 文件里。该文件通过opencompass/humaneval_pro这一 HuggingFace 数据集路径加载配置见 humaneval_pro_gen_3dc067.py。官方评测结果OC results仓库文档 README.md 中公布了 OpenCompass 评测OC下三个模型的 pass1百分制结果模型pass1qwen2.5-coder-7b-instruct-hf65qwen2.5-14b-instruct-hf67deepseek-v2-lite-chat-hf35其中qwen2.5-14b-instruct-hf在该组对比中取得最高分67deepseek-v2-lite-chat-hf得分最低35。这些结果可作为你在本机复现评测时的参考基准用于校验配置、后处理与评估链路是否工作正常。CodeEval-pro 对照结果同一份文档还给出了同一批模型在CodeEval-pro评测环境下的对照结果模型pass1qwen2.5-coder-7b-instruct-hf65qwen2.5-14b-instruct-hf65deepseek-v2-lite-chat-hf28对比两组数据可以发现OC 与 CodeEval-pro 环境下的得分存在差异例如qwen2.5-14b-instruct-hf分别为 67 与 65说明评测后端/执行环境会影响最终 pass1 数值。因此在进行跨环境结果比较时应明确标注评测环境避免直接横向对比不同管道产出的分数。评测配置详解HumanEval Pro 的正式评测配置位于 humaneval_pro_gen_3dc067.py它沿用了 OpenCompass 数据集配置的经典三段式结构reader_cfg、infer_cfg、eval_cfg。reader_cfg输入输出列定义humanevalpro_reader_cfg dict( input_columns[raw_problem, new_problem], output_columntest_code)input_columns喂给模型的输入列即raw_problem原始问题与new_problem依赖第一个解法的新问题output_column输出列test_code即用于校验答案的测试代码供评估器读取。infer_cfg提示词模板与推理器humanevalpro_infer_cfg dict( prompt_templatedict( typePromptTemplate, templatedict(round[ dict( roleHUMAN, promptPROMPT_WRAPPER), ])), retrieverdict(typeZeroRetriever), inferencerdict(typeGenInferencer))PROMPT_WRAPPER是内置的提示词模板要求模型把两个问题的解法统一放入一个 Python 代码块中且代码块内不含任何无关内容PROMPT_WRAPPER You are an exceptionally intelligent coding assistant that consistently delivers accurate and reliable responses to user instructions. Write a solution of python file to the following problems, the solution of the second problem requires single or multiple calls to the first solution. python {raw_problem} {new_problem}Please put the two solutions within the Python code block provided below, and make sure that the block contains no other unrelated content:- ZeroRetriever零样本检索不做示例选取直接使用完整提示词 - GenInferencer生成式推理器适用于代码生成这类文本续写任务。 ### eval_cfg远程评估服务 python humanevalpro_eval_cfg dict( evaluatordict(typeHumanevalProEvaluator, ip_addresshttps://opencompass-multiple-evaluator.hf.space) )typeHumanevalProEvaluatorHumanEval Pro 专属评估器实现见 humaneval_pro.pyip_address远程代码执行服务的地址。评测时评估器会把模型生成的代码与测试代码拼接后发送到该服务执行并根据返回结果统计 pass1。这意味着运行该评测需要具备访问该远程服务的网络条件。数据集注册humanevalpro_datasets [ dict( abbrhumaneval_pro, typeHumanevalevalProDataset, pathopencompass/humaneval_pro, reader_cfghumanevalpro_reader_cfg, infer_cfghumanevalpro_infer_cfg, eval_cfghumanevalpro_eval_cfg,) ]abbr评测结果中使用的简称humaneval_protype数据集类HumanevalevalProDataset其load方法读取 JSON 并转换为Dataset见 humaneval_pro.pypath数据在 HuggingFace 上的仓库路径。而 humaneval_pro_gen.py 作为正式入口配置通过read_base()直接继承上述humanevalpro_datasets保持配置的单一来源from mmengine.config import read_base with read_base(): from .humaneval_pro_gen_3dc067 import humanevalpro_datasets # noqa: F401, F403采样变体repeat_gen 配置仓库还提供了采样变体 humaneval_pro_repeat_gen_3dc067.py与正式版唯一的差异是数据集注册中增加了两个字段humanevalpro_datasets [ dict( abbrhumaneval_pro, typeHumanevalevalProDataset, pathopencompass/humaneval_pro, reader_cfghumanevalpro_reader_cfg, infer_cfghumanevalpro_infer_cfg, eval_cfghumanevalpro_eval_cfg, n5, k3) ]n5, k3表示每个问题采样 5 次生成、取前 3 次通过率passk 计算所需用于需要多次采样统计通过率的实验场景。其余 reader / infer / eval 三段配置与正式版完全一致。数据集实现与评估链路数据加载HumanevalevalProDataset.loadhumaneval_pro.py先通过get_data_path解析本地或远程路径随后读取 JSON 文件将每条记录依次追加进列表并转换为datasets.Datasetstaticmethod def load(path, local_modeFalse): path get_data_path(path, local_modelocal_mode) dataset [] with open(path, encodingutf-8) as f: raw_data json.load(f) for data in raw_data: dataset.append(data) return Dataset.from_list(dataset)评估器评分逻辑HumanevalProEvaluator继承自CodeEvaluatorcode_evaluator.py其score方法humaneval_pro.py执行三步组装测试用例对每条样本用_process_completions提取模型输出中的代码拼接测试代码test_code组成包含name、language、code的字典同时用PROMPT_WRAPPER生成对应的提示词记录批量送评调用_evaluate将全部测试用例一次性发送到远程评估服务统计结果根据服务返回的status OK判定通过计算pass1 100 * correct / total_count并附带逐条details。底层 CodeEvaluator 的行为CodeEvaluator的几个关键机制决定了评测的可靠性代码提取_extract_code通过正则r\w*\n(.*?)提取 Markdown 代码块内容优先取第一个代码块这与数据集通用后处理函数humaneval_postprocess_v2humaneval.py逻辑一致远程执行_code_eval_service基于gradio_client调用评估服务的/evaluate接口支持 dict / list / 文件路径三种输入超时会被捕获并转为失败状态重试机制_evaluate默认retry5连接失败时每 30 秒重试一次直至成功或耗尽重试次数避免临时网络抖动导致整批评测失败结果汇总_process_results将每个用例的原始输出回填prompt字段统计正确用例数并返回pass1与details。如何复现文档中的评测结果在 dataset-index.yml 中humaneval_pro的配置路径被注册为opencompass/configs/datasets/humaneval_pro/humaneval_pro_gen.py因此可以通过 OpenCompass 的run.py直接启动评测。基本流程如下# 1. 克隆并安装 OpenCompass 后使用内置配置启动评测 python run.py --models model_config \ --datasets humaneval_pro \ --work-dir outputs/humaneval_pro其中model_config为你要评测的模型配置文件例如文档结果表中对应的 qwen2.5 / deepseek-v2 系列模型。需要说明的适用前提评测依赖远程代码执行服务https://opencompass-multiple-evaluator.hf.space请确保运行环境可以访问该地址否则评估阶段会因连接失败而报错复现n5, k3的多采样版本可在数据集中显式传入对应配置或参考 humaneval_pro_repeat_gen_3dc067.py 编写自定义数据集配置最终分数会受模型版本、采样参数、评估服务状态影响文档中的数值是特定时间点的快照建议以本机复现结果为准。此外仓库中的 examples/eval_codebench_full.py 也引用了 humaneval_pro 相关配置可作为组合多个代码基准批量评测的参考样例。小结HumanEval Pro 在 OpenCompass 中的接入路径清晰配置层humaneval_pro_gen.py → humaneval_pro_gen_3dc067.py负责声明数据、提示词与评估器实现层humaneval_pro.py 与 code_evaluator.py负责加载数据、拼接测试代码并远程执行校验最终产出 pass1。理解了这三层你不仅能复现文档中的评测结果也能基于同一套骨架扩展出自己的多文件代码生成基准评测。赞分享模型评测人工智能大模型AI 评测【免费下载链接】opencompassOpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, etc, across 100 datasets covering knowledge, reasoning, coding, science, language, long-context, and safety.项目地址https://gitcode.com/gh_mirrors/op/opencompass点击查看免费下载相关推荐OpenCompass 代码评测实战基于 HumanEval 与 MBPP 的 pass1 / passk 完整配置指南OpenCompass 代码评测实战基于 HumanEval 与 MBPP 的 pass1 / passk 完整配置指南 导读 本文以 humaneval模型评测人工智能大模型AI 评测OpenCompass 代码评测服务实战指南基于 Docker 隔离环境完成 HumanEval-X 与 DS1000 评测OpenCompass 代码评测服务实战指南基于 Docker 隔离环境完成 HumanEval X 与 DS1000 评测 导读 LLM 生成的代码可能存在模型评测人工智能大模型AI 评测AI去味化的未来33种模式之后Humanizer的路线图与社区动态AI去味化的未来33种模式之后Humanizer的路线图与社区动态 Humanizer 是一款面向 AI 写作的「AI去味化」工具它以纯 Markdown模型评测人工智能大模型AI 评测上一篇MAX Python 请求类型全解析max.pipelines.request 模块架构与 OpenResponses API 实战指南下一篇Orleans Grain 接口版本化与向后兼容滚动升级中的契约安全指南创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
02
RELATED NEWS

相关资讯

更多网站建设与数字化升级内容

03
WHY YAOTU

想打造同款高转化官网?

懂行业、懂生意,从建站到增长一站式陪跑

◈

场景化定制

不做模板站,围绕你的业务场景量身设计,小众不撞款。

◐

营销型架构

以转化目标组织内容与路径,让官网真正带来询盘。

▲

全周期服务

设计、开发、运营、运维一体,上线只是开始。

免费获取你的建站方案

留下需求,专属顾问 24 小时内为你输出方案建议。