ClimateEval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change
Summary
Presents unified benchmark aggregating climate-related NLP evaluation datasets. Introduces Guardian Climate News Corpus with 6 climate-related categories. Includes 13 datasets covering stance detection, claim verification, news classification. Evaluates open-access baseline LLMs (2B-70B parameters) finding models struggle with domain-specific entities and seemingly simple tasks.
Themes
Keywords
NLP benchmark, climate change, text classification, claim verification, LLM evaluation
Poster
Click image to open full size