This file was generated on 2026-03-23 by LUKAS KORNHER A GENERAL INFORMATION 1. Title of the dataset: "Replication Data for: Baumüller and Kornher (2024). Inside the crowd: Assessing the suitability of SMS-based surveys to monitor the food security situation in Uganda" 2. Brief description of the research project and its aims: SMS-enabled surveys are gaining traction as a rapid, low-cost means of monitoring food security situations for early warning systems. However, such surveys run the risk of yielding biased results, for instance due to selection, attrition or non-response bias. . To assess the suitability of SMS-enabled surveys for food security monitoring in light of these potential biases, we conducted monthly surveys of 2000 respondents across Uganda for one year. A filtering approach was used to ensure a representative sample. We evaluate the likely accuracy of the data by triangulating the responses with high-frequency data from our own face-to-face household surveys as well as externally collected phone survey data. The analysis suggests that SMS-based surveys can be a promising tool to measure changes in food security status over time but perform less well with regard to measuring the actual food security status. Responses related to the general food situation (rather than dietary diversity, food consumption or market prices) emerged as the most reliable indicator. To investigate implications of selection bias on the results, we use different scenarios with variations in sample composition and size. Even biased samples, e.g. in terms of gender, location or age, show comparable trends, but a minimum sample size is required to obtain accurate results. 3. Author Information A. Investigator Contact Information Name: Dr Lukas Kornher Institution: German Institute of Development and Sustainability (IDOS) Address: Tulpenfeld 6, 53111 Bonn Email: lkornher@uni-bonn.de/lukas.kornher@idos-research.de 4. Date of data collection: 2020-02 to 2021-04 5. Information about funding sources that supported the collection of the data: German Federal Ministry for Economic Cooperation and Development: 2014.0689.1 German Federal Ministry for Economic Cooperation and Development: 2014.0690.9 6. Language of the dataset: English 7. Geographic location of data collection: Uganda: Central Region (without Kampala and Wakiso) Region 1: Buikwe, Bukomansimbi, Butambala, Buvuma, Gomba, Kalangala, Kalungu, Kayunga, Kiboga, Kyankwanzi, Luweero, Lwengo, Lyantonde, Masaka, Mityana, Mpigi, Mubende, Mukono, Nakaseke, Nakasongola, Rakai, Ssembabule Districts Eastern Region 2: Bugiri, Bukedea, Busia, Butaleja, Buyende, Iganga, Jinja, Kaliro, Kamuli, Kibuku, Luuka, Mayuge, Namayingo, Namutumba, Pallisa, Tororo Districts Northern-West Region 3: Adjumani, Amuru, Arua, Gulu, Koboko, Maracha, Moyo, Nebbi, Nwoya, Yumbe, Zombo Districts Karamoja/Mount Elgon Region 4: Abim, Agago, Bududa, Amudat, Bulambuli, Bukwa, Kaabong, Kitgum, Kapchorwa, Kotido, Lamwo, Kween, Moroto, Manafwa, Nakapiripirit, Napak, Mbale, Pader, Sironko Districts Northern-East Region 5: Amuria, Budaka, Alebtong, Amolatar, Apac, Dokolo, Kaberamaido, Kole, Katakwi, Lira, Kumi, Soroti, Serere, Ngora, Otuke, Oyam Districts Western Region 6: Buliisa, Bundibugyo, Bushenyi, Hoima, Kabarole, Kamwenge, Kasese, Kibaale, Kiryandongo, Kyegegwa, Kyenjojo, Masindi, Ntoroko, Rubirizi Districts Western-South Region 7: Buhweju, Ibanda, Isingiro, Kabale, Kanungu, Kiruhura, Kisoro, Mbarara, Mitooma, Ntungamo, Rukungiri, Sheema Districts B DATA & FILE OVERVIEW 1. File List: 1) KornherBaumueller2024.dta This dta files contains primary data from the SMS survey and High Frequency Panel Survey (HFPS) conducted by the authors. It includes the following variables: (More details can be found in the codebook.txt file) Survey_Comple~e double %tc Survey_Completed_time *Age float %10.0g Age Age_Group float %10.0g Age_Group BirthYear int %10.0g BirthYear Gender double %10.0g Gender Region float %20.0g Region Region Urban_Rural str19 %19s Urban_Rural HouseHold byte %10.0g Household size Farming str3 %9s Farming FoodSituation double %10.0g FoodSituation FoodSituation FoodConsumption double %30.0g FoodConsumption FoodConsumption MarketFood double %22.0g MarketFood MarketFood FoodPrice double %22.0g FoodPrice FoodPrice HaveProblems double %10.0g HaveProblems HaveProblems wave float %9.0g Wave Number SMS Survey UserId str21 %21s User Id SMS Survey district int %10.0g District number of respondent region_code float %9.0g region Region code of regions included in SMS Survey dd_score float %9.0g Dietary Diversity Score common_wave float %12.0g common_wave Wave Number of Common wave SMS and HFPS common_wave_wb float %12.0g common_wave_wb Wave Number of Common wave SMS and HFTS data str4 %9s Variable indicating source of data SMS vs. HFPS region_included float %9.0g Dummy=1 if region included in both SMS and HFPS time_included float %9.0g Dummy=1 if time period included in both SMS and HFPS FoodSitua~_poor float %9.0g Dummy=1 if Food Situtation is poor or very poor FoodSitua~vpoor float %9.0g Dummy=1 if Food Situtation is very poor HouseHold2 float %9.0g Household size squared *Age2 float %9.0g Age squared data2 str3 %9s Variable indicating source of data SMS vs. HFTS agriculture float %9.0g Dummy=1 if household lives in farming household urban float %9.0g Dummy=1 if respondent lives in urban area *whenever the age of the household head in the HFPS was not available it was replaced by the average age of the remaining household members (which explains decimal values in Age and Age2) The variables data and data2 contain information from which source the specific observation comes from. Variables include self-explanatory labels. 2) replication.do This do file contain all manipulations and calculations on the data to obtain the results presented in the paper. In the do file, it is indicated which table and figure is generated from the output. 3) 5 files random_draw** Some of the analysis the paper is based on a random draw of “x” respondents from the entire data. Re-rerunning the random draws will generate slightly different results for Figure 4. Therefore, we also provide the stata files of the random draws used to generate Figure 4. These files are: "random_draw 1.dta" 1% (equivalent of 20 respondents per wave) "random_draw 2,5 .dta" 2.5% (equivalent of 50 respondents per wave) "random_draw 10.dta" 10% (equivalent of 200 respondents per wave) "random_draw 25 .dta" 25% (equivalent of 500 respondents per wave) "random_draw 50 .dta" 50% (equivalent of 1000 respondents per wave) 4) HFTS_clean.do This do file generates the dataset “hfwb.dta” from the raw data files of the Uganda High-Frequency Phone Survey 2020-2023. These data can be downloaded on https://microdata.worldbank.org/index.php/catalog/3765/get-microdata up on registration. 2. Are there multiple versions of the dataset? no 3. Relationship between files: (1) and (2) are generated from (3) via (4) 4. Additional related data collected that was not included in the current data package: n.a. C SHARING/ACCESS INFORMATION 1. Was data derived from another source? Uganda Bureau of Statistics (UBOS). (2020). High-Frequency Phone Survey 2020-2024 [Data set]. World Bank, Development Data Group. https://doi.org/10.48529/PG7V-3896 These data can be downloaded on https://microdata.worldbank.org/index.php/catalog/3765/get-microdata up on registration. 2. Licenses/restrictions placed on the data: Licensed under CC0 1.0 3. Links to publications that cite or use the data: D METHODOLOGICAL INFORMATION 1. Description of methods used for collection/generation of data: The survey was implemented by the private company GeoPoll through its 2-way SMS system. Their SMS API enables programmatically sending and receiving messages through short codes that are free to end users, allowing them to respond even if they do not have airtime. The survey was scripted in GeoPoll’s own platform. Upon completion of each survey round, the GeoPoll team reviewed the data, checking for unusual survey response patterns, drop-offs on specific questions or potential skews in the data. 2. Methods for processing the data: Submitted data was generated from individual csv files. Labels were added. No further processing was done. 3. Instrument- and/or software-specific information needed to interpret the data: Stata 16.0 4. People involved in sample collection, processing, analysis and/or submission: The survey was implemented by the private company GeoPoll through its 2-way SMS system. Their SMS API enables programmatically sending and receiving messages through short codes that are free to end users, allowing them to respond even if they do not have airtime. The survey was scripted in GeoPoll’s own platform. Upon completion of each survey round, the GeoPoll team reviewed the data, checking for unusual survey response patterns, drop-offs on specific questions or potential skews in the data. 5. Describe any quality-assurance procedures performed on the data: n.a. 6. Standards and calibration information: n.a.