## Multilingual Eval Suite This is going to be based on Maria's report. Here is a list based on what's implemented so far in [oellm-eval](https://github.com/OpenEuroLLM/oellm-eval) - [x] Belebele - [x] XWinograd - [x] Global MMLU - [x] INCLUDE - [x] XStoryCloze - [x] MGSM - [x] GPQA - [x] ARC challenge_mt - [x] SIB-200 - [x] Global PIQA - [x] [Open Subtitles](https://huggingface.co/datasets/Helsinki-NLP/OpenSubtitles2024-40-langs-15-movies) - #70 - [x] [FLORES-200](Muennighoff/flores200) - #101 - [ ] [GLOBAL MGSM](https://huggingface.co/datasets/CohereLabs/global-mgsm) - #103 - [ ] [XCopa](https://huggingface.co/datasets/cambridgeltl/xcopa) - #66 - [ ] [Hellaswag](https://huggingface.co/datasets/alexandrainst/m_hellaswag) - #102 - [x] [PolyMath](https://huggingface.co/datasets/Qwen/PolyMath) - #94 - [ ] [MMMLU](https://huggingface.co/datasets/openai/MMMLU) - #98 - [x] [MMLU-ProX](https://huggingface.co/datasets/li-lab/MMLU-ProX) - #99 - [ ] [MultiBlimp](https://huggingface.co/datasets/jumelet/multiblimp) - #65 - [x] [xcsqa]() - #100 - [ ] [MathNet](https://huggingface.co/datasets/ShadenA/MathNet) - Note: Contains Images as part of the prompt - [ ] PolyMath-TUM - [x] XNLI - #105 - [x] PAWS-x - #106 - [ ] Swesat - [ ] WikiAnn NER / PAN-X - [ ] Tatoeba challenge - [ ] Taxi1500 - [ ] MENLO - [ ] Doclevel-mt - [ ] BFCL - [ ] tau2-bench - [ ] [BLeND](https://huggingface.co/datasets/nayeon212/BLEnD) - [ ] [Toxigen](https://huggingface.co/datasets/toxigen/toxigen-data)
Multilingual Eval Suite
This is going to be based on Maria's report. Here is a list based on what's implemented so far in oellm-eval