Leaderboard
A Maltese Evaluation Language Benchmark 🇲🇹
⌛ Loading leaderboard data from the MLRS datasets…
- "headers": [
- "T",
- "Model",
- "N-Shot",
- "Version",
- "Average (All) ⬆️",
- "Average (NLU) 🧠",
- "Average (NLG) ✍️",
- "Sentiment Analysis (F1)",
- "SIB200 (F1)",
- "Taxi1500 (F1)",
- "Maltese News Categories (F1)",
- "MultiEURLEX (F1)",
- "Belebele (Accuracy)",
- "OPUS-100 EN→MT (BLEU)",
- "OPUS-100 EN→MT (ChrF)",
- "Flores-200 EN→MT (BLEU)",
- "Flores-200 EN→MT (ChrF)",
- "WebNLG (ChrF)",
- "WebNLG (Rouge-L)",
- "EUR-Lex-Sum (ChrF)",
- "EUR-Lex-Sum (Rouge-L)",
- "Maltese News Headlines (ChrF)",
- "Maltese News Headlines (Rouge-L)",
- "Type",
- "Maltese Training",
- "#Languages",
- "Architecture",
- "Precision",
- "Hub License",
- "#Params (B)",
- "Hub ❤️",
- "Available on the hub",
- "Model SHA"
- "data": [
- []
- "metadata": null
- "headers": [
- "T",
- "Model",
- "N-Shot",
- "Version",
- "Average (All) ⬆️",
- "Average (NLU) 🧠",
- "Average (NLG) ✍️",
- "Sentiment Analysis (F1)",
- "SIB200 (F1)",
- "Taxi1500 (F1)",
- "Maltese News Categories (F1)",
- "MultiEURLEX (F1)",
- "Belebele (Accuracy)",
- "OPUS-100 EN→MT (BLEU)",
- "OPUS-100 EN→MT (ChrF)",
- "Flores-200 EN→MT (BLEU)",
- "Flores-200 EN→MT (ChrF)",
- "WebNLG (ChrF)",
- "WebNLG (Rouge-L)",
- "EUR-Lex-Sum (ChrF)",
- "EUR-Lex-Sum (Rouge-L)",
- "Maltese News Headlines (ChrF)",
- "Maltese News Headlines (Rouge-L)",
- "Type",
- "Maltese Training",
- "#Languages",
- "Architecture",
- "Precision",
- "Hub License",
- "#Params (B)",
- "Hub ❤️",
- "Available on the hub",
- "Model SHA"
- "data": [
- []
- "metadata": null
MELABench evaluates language model capabilities on Maltese. Currently, the following tasks are supported:
NLU:
NLG:
The leaderboard is developed and maintained by people managing MLRS. We plan to expand our initial work with more tasks, if you would like to contribute your data, please reach out! If you would like to include results for models/setups we did not include, we also accept submissions.
This work was introduced in MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP.
To include new results on this benchmark, follow the instructions on our GitHub Repository. You can then upload the output files which should include the configuration/results file and all the prediction files. In addition, we ask for additional metadata about model training.
model | revision | precision | n_shot | prompt_version | seed | status |
|---|---|---|---|---|---|---|
model | revision | precision | n_shot | prompt_version | seed | status |
|---|---|---|---|---|---|---|
model | revision | precision | n_shot | prompt_version | seed | status |
|---|
model | revision | precision | n_shot | prompt_version | seed | status |
|---|---|---|---|---|---|---|
model | revision | precision | n_shot | prompt_version | seed | status |
|---|
✉️✨ Submit your model here!
How to model is trained.
The last stage of training in which Maltese was included.