This file was generated on 07-07-2025 by Marion Pitz A GENERAL INFORMATION 1. Title of the dataset: Data from dissertation "Transcriptomic regulation of hybrid vigor in immortalized backcross populations of maize (Zea mays L.) 2. Brief description of the research project and its aims: "This thesis aimed to close these knowledge gaps and advance the molecular understanding of heterosis. We analyzed 112 lines of the intermated B73xMo17 recombinant inbred line (IBM-RIL) population of maize and their backcrosses to B73 and Mo17. These backcross hybrids contain heterozygous and homozygous genomic regions and allow for the identification of genomic locations regulating specific gene expression patterns." 3. Author Information A. Investigator Contact Information Name: Marion Pitz Institution: University of Bonn Address: Römerstraße 164 Email: mpitz@uni-bonn.de B. Project Supervisor (Principal Investigator) Contact Information Name: Prof. Dr. Frank Hochholdinger Institution: University of Bonn Address: Römerstraße 164 Email: hochhold@uni-bonn C. In case of questions related to this dataset, please contact: Name: Marion Pitz Institution: University of Bonn Address: Römerstraße 164 Email: mpitz@uni-bonn.de 4. Date of data collection : 2021-06-01 - 2024-12-14 5. Information about funding sources that supported the collection of the data: This work was supported by the Deutsche Forschungsgemeinschaft (DFG) Next Generation Sequencing Competence Network (NGS-CN; project 423957469) grant HO 2249/18-1 to FH. 6. Language of the dataset: English 7. Geographic location of data collection : Bonn, Germany B DATA & FILE OVERVIEW 1. File List: Supplementary data files for Chapter 3 in the digital appendix Data_S1_exclusion_table.xlsx Details of retained and excluded samples. Data_S2_Loci_listB73vsMo17.xlsx List and details of identified loci with different alleles in B73 and Mo17 samples. Data_S3_IBM-RIL_haplotype_regions.xlsx Specifications of start and end position of each identified Mo17, B73 or IBM-RIL specific region in each IBM-RIL. Data_S4_eQTL_details.xlsx Details on identified eQTL Data_S5_MPH_pheno.xlsx Phenotypic mid-parent heterosis values of each hybrid. Data_S6_pattern_ref.xlsx SPE pattern of all genes in reference hybrids and their eQTL. Data_S7_TWAS_genes.xlsx All genes and details identified in the TWAS analysis. Data_S8_TSG.xlsx Details on SPE pattern across hybrids of TWAS candidate genes. Data_S9_Zm00001eb339600_B73.txt BLAST result for candidate gene. Supplementary data files for Chapter 4 in the digital appendix Data_S1_NAG_refs.xlsx Non-additive pattern of all genes in reference hybrids. Data_S2_overview.xlsx Summarised details of additive and non-additive genes in all hybrids. Data_S3_het_prop.xlsx Proportions of heterozygous regions in the hybrids. 2. Are there multiple versions of the dataset? no 4. Additional related data collected that was not included in the current data package: Raw transcriptome data in NCBI Bioproject ID PRJNA923128 (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA923128). C SHARING/ACCESS INFORMATION 1. Was data derived from another source? : No 2. Licenses/restrictions placed on the data: This document is licensed under the Creative Commons CC-BY-4.0 International License. 3. Links to publications that cite or use the data : https://doi.org/10.1101/2024.10.30.620956, https://doi.org/10.1111/nph.70128 4. Links to other publicly accessible locations of the data : https://doi.org/10.1101/2024.10.30.620956, https://doi.org/10.1111/nph.70128 5. Links/relationships to ancillary datasets: NCBI Bioproject ID PRJNA923128 D METHODOLOGICAL INFORMATION 1. Description of methods used for collection/generation of data: Documented in material and methods of https://doi.org/10.1101/2024.10.30.620956 2. Methods for processing the data: We studied the phenotypic and transcriptomic plasticity of the IBM-RIL backcross hybrids relative to their parental inbred lines. A selection of 112 IBM-RILs were used as paternal inbred lines, corresponding to 112 B73×IBM-RIL and 112 Mo17×IBM-RIL backcross hybrids (Figure 1A). To optimally fit the experimental design and increase the precision of subsequent pairwise comparisons with the common parental inbred lines and reference hybrids both common maternal inbred lines B73 and Mo17 and the two reference hybrids B73×Mo17 or Mo17×B73, respectively, were included (Figure 1B). For each sample, 25 kernels of the same genotype (parent or hybrid) were surface sterilized in 10% H2O2 for 20 min, rinsed with distilled water and afterwards pre-germinated in filter paper rolls with five kernels each in a climate chamber with a 16 h light (26 °C), 8 h dark (21 °C) cycle in distilled water11. After three days, eight seedlings per genotype with approximately the same length of primary root and, if already present, shoot length were selected and transferred into a row of an aeroponic growth system for four additional days. Each aeroponic growth system (“Elite Klone Machine 96”, TurboKlone, USA) was composed of 12 rows each with eight planting sites. Thus, we could fit 12 different genotypes into one aeroponic growth system and eight systems at the same time into our climate chamber (16 h light, 26 °C; 8 h dark, 21 °C) (Figure S7). We analyzed three independent biological replicates, each comprising all IBM-RIL inbred lines and hybrids. Due to the large number of samples and space limitations, the different genotypes of each biological replicate were grown in four batches distributed across four weeks, also called alpha-design with incomplete blocks30. Within each batch, the eight aeroponic growth systems were randomly assigned to eight positions in the climate chamber. Three successive rows of an aeroponic growth system represent one triplet. To each triplet an IBM-RIL and its corresponding B73 and Mo17 backcross hybrids, or both common maternal inbred lines B73 and Mo17 and one of the two reference hybrids B73×Mo17 or Mo17×B73, respectively were assigned. The randomization process was conducted at the replicate level for the triplets, whereas it was ensured, that in each batch two reference triplets each of B73, Mo17, B73xMo17 and B73, Mo17, Mo17xB73 were distributed. Thus, in each batch we surveyed 30 IBM-RIL triplets and both reference hybrids and the common inbred lines B73 and Mo17 in two additional triplets (Figure S7). So that in total, 3 samples of each IBM-RIL and each backcross hybrid, 48 samples (biological replicates) of the maternal inbred lines B73 and Mo17, and 24 samples of the reciprocal hybrids B73×Mo17 and Mo17×B73 as reference hybrids were analyses. In other words, each independent biological replicate contained 384 samples (in total: 384 x 3 replicates = 1,152 samples). The number of individual samples per replicate was designed to also fit one sequencing run on the NovaSeq 6000 S4 flow cell machine (Illumina, San Diego, USA), described later. Seven days after germination, all seedlings per sample (maximum eight seedlings) were removed from the aeroponic growth system and the seedling root system was scanned using an Epson Expression 12000XL scanner (Epson, Meerbusch, Germany) with up to four plants per image. The resulting images were cropped to create single plant images that only showed the root system and the maize kernel. We used the RootPainter software client (version 0.1.0) and server component (version 0.2.7) to train a convolutional neural network to recognize and segment roots in images31. We then analyzed the segmented images in a batch using RhizoVision Explorer (version 2.0.3)32. After inspection and cleanup, we determined the total root length, total root volume and number of root tips for each plant for subsequent analysis (Details in Supplement Material SM1). After imaging the seedlings, the primary root was separated from the kernel to collect (i) the proximal first centimeter with emerged lateral roots in 80% Ethanol to count the number of lateral roots per cm as density and (ii) the distal region of the primary root, composed of the root tip and the meristematic zone followed by the elongation zone, in liquid nitrogen for subsequent RNA extraction. The filtered SNP data were used to classify each IBM-RIL genome into B73 or Mo17 regions and to mask regions which were not B73 or Mo17. A distance-function was used to calculate the distance between the IBM-RIL specific loci. Loci with a distance of <2.5 Mbp were grouped together as a block. Blocks containing a minimum of 10 IBM-RIL-specific loci, with ≥5 of those being homozygous, were identified and masked as IBM-RIL-specific third origin regions. The start and end positions of these regions were recorded and loci within those regions were dropped. Next, a sliding window approach was used to eliminate singular loci that did not match their surrounding loci. A window of 15 loci was used, and ≥11 had to be homozygous for the Mo17 allele for the window to be considered a Mo17 window. For a B73 window, ≥12 out of 15 loci must have homozygous B73 alleles. The values for the windows were obtained by computing the minimum number of matching loci in a 15-loci window across B73 and Mo17 samples. Otherwise, the window was considered ambiguous40. Loci within an ambiguous window were dropped, as well as loci which were classified differently from their window. The previously mentioned distance-function was utilised to calculate the distance between the remaining loci. Loci that carried the same allele and which were less than 0.5 Mbp apart were grouped together as a block, and all blocks were retained. The start and end positions of these blocks were recorded as the Mo17 and B73 regions within each IBM-RIL (Data file S9). Two IBM-RILs which had more than 50% of their genomes consisting of IBM-RIL specific regions from a third parental origin were excluded along with their hybrids (Data file S1), leaving 834 samples for final analyses. The data set of each triplet reported by the HaplotypeCaller was filtered to only include loci within the B73 or Mo17 regions of the IBM-RILs and within exons of protein-coding genes. We checked for all protein-coding genes whether they were located in a B73 or Mo17 region or masked as neither a B73 or Mo17 region, or whether they were located in a genomic region without SNP information. This verification was performed for each IBM-RIL separately. Centromere locations of the 10 chromosomes were taken from the genome assembly of MaizeGDB by selecting the “Knobs, centromeres and telomeres” information https://jbrowse.maizegdb.org41. The proportion of heterozygous to homozygous regions was calculated for each backcross hybrid by dividing the total lengths of classified heterozygous regions (B73 regions of the IBM-RIL for the Mo17×IBM-RIL and Mo17 regions of the IBM-RIL for B73×IBM-RIL) by the total lengths of all classified regions (not considering IBM-RIL specific masked regions and regions without SNP information). Data filtering, Alignment, Activity and expression analyses, eQTL analyses and TWAS described in https://doi.org/10.1101/2024.10.30.620956 and https://doi.org/10.1111/nph.70128 Files are named after their reference in the manuscript/thesis where their occur (S1...) followed by a short description. E DATA-SPECIFIC INFORMATION FOR: Data_S4_eQTL_details.xlsx 1. Variable list including full names and definitions (please spell out abbreviated words) of column headings for tabular data: eQTL: expression Quantitative Trait Locus, LOD: Logarithm of Odds, FDR: False Discovery Rate, Conf-intervall_low: Lower end of the confidence intervall around the eQTL, Conf-intervall_high: Upper end of the confidence intervall around the eQTL 2. Units of measurement used: Mega base pairs 3. Missing data codes/symbols: NA