This file was generated on 2026-07-23 by Pauline Suski A GENERAL INFORMATION 1. Title of the dataset: Data for: Legume Products and Food Lifestyles: Data from a Comprehensive Consumer Survey 2. Brief description of the research project and its aims: This dataset contains results from a survey (n = 1,998) on German diets with a special focus on legume consumption and how this is associated with food-related lifestyles. This includes food consumption, food-related lifestyles, sociodemographics, legume consumption as well as knowledge, associations and opinions of legumes and legume containing meals. Data was collected between December 2024 and January 2025. 769 variables can be found in the dataset. The dataset is part of the StrahL project to develop target group specific strategies for an increased consumption of legumes in Germany. The dataset was used in the project to identify and describe various target groups (segments) and their legume consumption. 3. Author Information A. Investigator Contact Information Name: Pauline Suski Institution: University of Bonn - Institute for Food and Resource Economics Address: Am Probsthof 49, Bonn Email: pauline.suski@ilr.uni-bonn.de Name: Fabienne Erben Institution: Georg August University Göttingen Address: Kellnerweg 6, Göttingen Email: fabienne.erben@uni-goettingen.de Name: Daniel Mörlein Institution: Georg August University Göttingen Address: Kellnerweg 6, Göttingen Email: daniel.moerlein@uni-goettingen.de Name: Catherina Jansen Institution: Fulda University of Applied Sciences Address: Leipziger Straße 123, Fulda Email: catherina.jansen@oe.hs-fulda.de Name: Julian Quandt Institution: Corsus Research gUG Address: Großneumarkt 50, Hamburg Email: julian.quandt@uni-a.de Name: Anke Zühlsdorf Institution: Zuehlsdorf & Partner PartG Address: Philipp-Oldenbürger-Weg 27, Göttingen Email: azuehls@gwdg.de Name: Tonia Ruppenthal Institution: Fulda University of Applied Sciences Address: Leipziger Straße 123, Fulda Email: tonia.ruppenthal@oe.hs-fulda.de Name: Jana Rückert-John Institution: Fulda University of Applied Sciences Address: Leipziger Straße 123, Fulda Email: jana.rueckert-john@oe.hs-fulda.de Name: Dominic Lemken Institution: University of Bonn - Institute for Food and Resource Economics Address: Am Probsthof 49, Bonn Email: dominic.lemken@ilr.uni-bonn.de B. Project Supervisor (Principal Investigator) Contact Information Name: Pauline Suski Institution: University of Bonn - Institute for Food and Resource Economics Address: Am Probsthof 49, Bonn Email: pauline.suski@ilr.uni-bonn.de C. In case of questions related to this dataset, please contact: Name: Dominic Lemken Institution: University of Bonn - Institute for Food and Resource Economics Address: Am Probsthof 49, Bonn Email: dominic.lemken@ilr.uni-bonn.de 4. Date of data collection: December 2024 to January 2025 5. Information about funding sources that supported the collection of the data: The work was supported by funds of the Federal Ministry of Agriculture, Food and Regional Identity (BMLEH) based on a decision of the Parliament of the Federal Republic of Germany via the Federal Office for Agriculture and Food (BLE) under the Protein Crop Strategy (Grant 2822EPS020). 6. Language of the dataset: English 7. Geographic location of data collection : Germany B DATA & FILE OVERVIEW 1. File List: survey_data-legumes_and_diets.sav (dataset in sav format for SPSS) survey_data-legumes_and_diets.csv (dataset in csv format) survey_data-legumes_and_diets.xlsx (dataset in xlsx format) List-Value-Labels.xlsx (Overview of Value labels of each variable) List-Variables.xlsx (Overview of Variables, including labels, types and measurment level) Survey_english.docx Codebook.docx (Overview of survey questions, response options, responses, variable names and types, branching logic, thematic blocks) 2. Are there multiple versions of the dataset? no 3. Relationship between files: Both List-files serve as additions to the datasets in .csv and .xlsx format as those formats do not include enough information on the variables (variable labels and value labels). The .sav dataset already includes all information on the variables and does not need the additional lists. The codebook grants a comprehensive overview for the dataset and is not intended to be directly used in statistical software, but to get familiar with the content and structure of the dataset in a readable format. Besides variables, labels and value labels, it contains actual responses for most variables for quick descriptive insights and also shows branching logic. Furthermore, the codebook provides additional background information on some variables, e.g. descriptions on the German school system to better understand response options, and starts with a content overview, based on 8 thematic blocks from the survey, as described in the corresponding data article. 4. Additional related data collected that was not included in the current data package: none C SHARING/ACCESS INFORMATION 1. Was data derived from another source? no 2. Licenses/restrictions placed on the data: CC BY 4.0 3. Links to publications that cite or use the data: Erben, F.; Suski, P.; Strack, M.; Lemken, D.; Mörlein, D. (2026): Consumer segmentation based on food-related lifestyles for targeted promotion of legume consumption in Germany. In: Food Research International 242, S. 119852. DOI: 10.1016/j.foodres.2026.119852. This manuscript is currently under review Suski, P., Erben, F., Mörlein, D., Jansen, C., Quandt, J., Zühlsdorf, A., Ruppentahl, T., Rückert-John, J., Lemken, D., under review. Legume Products and Food Lifestyles: Data from a Comprehensive Consumer Survey. Data in Brief. D METHODOLOGICAL INFORMATION 1. Description of methods used for collection/generation of data: The cross-sectional online survey took place in Germany and in German between December 2024 and January 2025, excluding the Christmas holidays. All participants were recruited through the panel provider’s platform and completed the survey remotely. Participation was voluntary and anonymous. A pre-test was conducted (50 respondents) for quality assessment and minor revisions were made to item wording and response categories. The median completion time was 21 minutes. Participants were recruited through quota sampling to ensure representation across age, gender, income, federal state, and education level. Participants received an incentive from the panel provider upon completion. 2 attention checks were included in the survey. The attached file Survey_english.docx contains all questions and answer options from the questionnaire, including the corresponding variables. It is an English translation of the original questionnaire that was German. 2. Methods for processing the data: Quality control measures was applied to identify and exclude invalid cases. The following criteria were used: • Quota screenouts (n = 2,285 excluded) • Failure to pass attention checks (n = 34 excluded) • Incomplete questionnaires (n = 138 excluded) • “Speeding” behaviour (completion time below 10.49 minutes (half the median); n= 80 excluded) • “Straightlining” in three content blocks (n = 85 excluded) • Inconsistent or contradictory response patterns (n = 38 excluded) A final sample of n= 1,998 complete and valid cases remained for analysis. Data cleaning was conducted using R version 4.4.1, while variable labeling and final formatting for further analysis were performed in SPSS v31.0.0, accommodating the software preferences of different data users. 3. Instrument- and/or software-specific information needed to interpret the data: The dataset with the fullest information is the file survey_data-legumes_and_diets.sav. .sav is a file format created and used in the commercial statistics software SPSS (created in SPSS v31.0.0). Other software might also be able to read the file, but this was not tested. The data files in .xlsx (for Microsoft Excel) and .csv (open format), also contain all variables and values, but not variable labels or value labels. Those are found in the additional files List-Variables.xlsx and List-Value_Labels.xlsx so that all data can be used independent of the specific software. 4. People involved in sample collection, processing, analysis and/or submission: Fabienne Erben took care of data collection and processing (quality control). Pauline Suski added labeling, final formatting and translation, as well as took care of submission and description. 5. Describe any quality-assurance procedures performed on the data: See above under point D.2 E DATA-SPECIFIC INFORMATION FOR: all files All information on the data (variables) are found in the attached documents: List-Variables.xlsx List-Value_Labels.xlsx Survey_english.docx Codebook.docx Five acronyms are relevant to know to read the data FFQ - Food frequency questionnaire FRL - Food Related Lifestyle eFRL - extended Food Related Lifestyle FNS - Food Neophobia Scale NA - not applicable Some cells in the dataset are filled with NA (not applicable) due to used forks. For examples, Q3 and Q4, which represent a double check for agreement in participation in case participants did not agree to participate in the first place (Q1), are not applicable for most participants. The food frequency questionnaire blocks were also developed with many forks in order to keep effort for participants low (e.g. Q98ff). In contrast, "none" refers to an actual given answer.