Jiayi Pan
917d96e06f
Fix doc error in evals ( #2654 )
2024-06-27 16:13:47 +00:00
Graham Neubig
cab7a288ca
Add NUM_WORKERS variable to run_infer.sh scripts for configurable woker settings ( #2597 )
...
* Add NUM_WORKERS variable to run_infer.sh scripts for configurable worker settings
* Update evaluation/webarena/scripts/run_infer.sh
---------
Co-authored-by: OpenDevin <opendevin@all-hands.dev >
2024-06-23 03:43:43 +00:00
Boxuan Li
feabc97aba
Evaluation time travel: build sandbox on the fly ( #2491 )
2024-06-20 20:22:02 -06:00
Boxuan Li
6f235937cf
Evaluation time travel: allow evaluation on a specific version ( #2356 )
...
* Time travel for evaluation
* Fix source script path
* Exit script if given version doesn't exist
* Exit on failure
* Update README
* Change scripts of all other benchmarks
* Modify README files
* Fix logic_reasoning README
2024-06-16 10:25:14 -04:00
RainRat
745ae42a72
fix typos ( #2352 )
2024-06-09 12:57:58 -07:00
yueqis
68d9ad61cf
Feat: Support Gorilla APIBench ( #2081 )
...
* removed unused files from gorilla
* Update run_infer.py, removed unused imports
* Update utils.py
* Update ast_eval_hf.py
* Update ast_eval_tf.py
* Update ast_eval_th.py
* Create README.md
* Update run_infer.py
* make lint
* Update run_infer.py
* fix lint
---------
Co-authored-by: yufansong <yufan@risingwave-labs.com >
2024-06-08 16:54:54 +00:00