mirror of https://github.com/All-Hands-AI/OpenHands.git synced 2026-04-29 03:00:45 -04:00

Files

Leo be251b11de Add AgentBench. (#2012 )

* Add AgentBench.

* Load the datasets from HF.

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Add helper functions.

* Add mock executor.

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Add retriv agent answer cmd.

* Adjust the dataset.
* Refine test results.

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Consolidate all AgentBench datasets and scripts into a single CSV dataset.

* Refactor dataset source.
* Update helper functions.

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Fix the CRLF problem.

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Separate the instance's workspace.

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Add cleanup logic and error handling for sandbox closure.

* Normalized dataset

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Update README.

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Update the prompt to capture the answer.

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Refactor script execution paths to use absolute container workspace path.

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Update AgentBench README.

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Delete useless functions.

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Update evaluation/agent_bench/README.md

* Add script to summarize test results from JSONL file in AgentBench

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Delete useless script and codes.

Signed-off-by: ifuryst <ifuryst@gmail.com>

* Update evaluation/agent_bench/scripts/summarise_results.py

---------

Signed-off-by: ifuryst <ifuryst@gmail.com>
Co-authored-by: Boxuan Li <liboxuan@connect.hku.hk>

2024-06-01 07:58:14 +00:00

agent_bench

Add AgentBench. (#2012 )

2024-06-01 07:58:14 +00:00

EDA

Support Entity-Deduction-Arena (EDA) Benchmark (#1931 )

2024-05-25 23:17:04 +08:00

gaia

update README for GAIA (#2054 )

2024-05-25 15:01:03 +00:00

humanevalfix

HumanEvalFix integration (#1908 )

2024-05-23 13:09:40 +00:00

logic_reasoning

Delete evaluation outputs files (#2152 )

2024-05-31 03:12:27 +00:00

mint

Add remaining subsets for MINT benchmark (#2142 )

2024-05-31 20:04:13 +00:00

regression

Feat: add stream output to exec_run (#1625 )

2024-05-16 14:37:49 +00:00

static

Add detailed tutorial for adding new evaluation benchmarks (#1827 )

2024-05-18 13:40:53 -04:00

swe_bench

SWE-bench: Add summarise utility script to view passed/failed task IDs (#2137 )

2024-05-31 12:32:17 +08:00

__init__.py

feat(SWE-Bench environment) integrate SWE-Bench sandbox (#1468 )

2024-05-15 16:15:55 +00:00

README.md

Support MINT benchmark (MATH, GSM8K subset) (#1955 )

2024-05-28 07:42:52 +00:00

TUTORIAL.md

docs: update tutorial docs (#1912 )

2024-05-20 14:40:31 +00:00

README.md

Evaluation

This folder contains code and resources to run experiments and evaluations.

Logistics

To better organize the evaluation folder, we should follow the rules below:

Each subfolder contains a specific benchmark or experiment. For example, evaluation/swe_bench should contain all the preprocessing/evaluation/analysis scripts.
Raw data and experimental records should not be stored within this repo.
- For model outputs, they should be stored at this huggingface space for visualization.
Important data files of manageable size and analysis scripts (e.g., jupyter notebooks) can be directly uploaded to this repo.

Supported Benchmarks

SWE-Bench: evaluation/swe_bench
HumanEvalFix: evaluation/humanevalfix
GAIA: evaluation/gaia
Entity deduction Arena (EDA): evaluation/EDA
MINT: evaluation/mint

Result Visualization

Check this huggingface space for visualization of existing experimental results.

Upload your results

You can start your own fork of our huggingface evaluation outputs and submit a PR of your evaluation results to our hosted huggingface repo via PR following the guide here.