AutoLearn loader
Authors
T. Eftimov, A. Gjorgjevikj, M. Martinc, G. Cenikj, D. Sasanski, R. Stojanov, S Dzeroski, B. Koroušić Seljak
Publication
Scientific Data, 2026
Abstract

General-purpose Large Language Models (LLMs) like Llama, GPT, and Mistral struggle with domain-specific challenges in food and nutrition, where data is fragmented, heterogeneous, and semantically complex. While fine-tuned LLMs have shown success in healthcare and life sciences, similar progress in food domains has been limited, largely due to the lack of high-quality, task-specific datasets. We present FoodBench, a curated benchmark dataset of question–answer pairs designed for training and evaluating LLMs in food and nutrition. It spans key tasks such as nutrient estimation, food traffic-light classification, synonym linking, cooking measurement conversion, and food named-entity recognition and linking. FoodBench enables robust performance evaluation across zero-, one-, and few-shot settings, laying the groundwork for trustworthy, domain-adapted language models. This resource supports advances in personalized nutrition, dietary assessment, and food system innovation. Evaluation of four general-purpose LLMs (Llama 3, Mistral, Gemma, Gemini) on FoodBench tasks shows limited performance across nutrient estimation, traffic-light classification, and food interoperability, even with few-shot prompting. These results highlight the need for domain-specialized LLMs fine-tuned on food data, while establishing FoodBench as a benchmark not only for assessing general-purpose models but also for guiding and evaluating fine-tuning efforts.