------------------------------------------------------------------------------------------------------------- This file was generated on 2025-07-22 by Samuel Nestor Meckoni and Boas Pucker A GENERAL INFORMATION 1. Title of the dataset: Genome sequence and annotation of Victoria cruziana 2. Brief description of the research project and its aims: The genome of a Victoria cruziana plant was sequenced with nanopore long reads. The genome sequence was assembled with Verkko2, scaffolding was conducted with CPhasing and the gene models were predicted by BRAKER3 and GeMoMa. The functional annotation was predicted based on sequence similarity to well characterized Arabidopsis thaliana sequences. 3. Author Information A. Investigator Contact Information Name: Melina Sophie Nowak Institution: TU Braunschweig Address: - Email: melina.nowak@tu-braunschweig.de B. Investigator Contact Information Name: Benjamin Harder Institution: TU Braunschweig Address: - Email: b.harder@tu-braunschweig.de C. Investigator Contact Information Name: Samuel Nestor Meckoni Institution: IZMB, University of Bonn Address: Kirschallee 1, 53115 Bonn, Germany Email: meckoni@uni-bonn.de D. Investigator Contact Information Name: Ronja Friedhoff Institution: TU Braunschweig Address: - Email: r.friedhoff@tu-bs.de E. Investigator Contact Information Name: Katharina Wolff Institution: IZMB, University of Bonn Address: Kirschallee 1, 53115 Bonn, Germany Email: k.wolff@uni-bonn.de F. Project Supervisor (Principal Investigator) Contact Information Name: Boas Pucker Institution: IZMB, University of Bonn Address: Kirschallee 1, 53115 Bonn, Germany Email: pucker@uni-bonn.de G. In case of questions related to this dataset, please contact: Name: Boas Pucker Institution: IZMB, University of Bonn Address: Kirschallee 1, 53115 Bonn, Germany Email: pucker@uni-bonn.de 4. Date of data collection: 2023-07-01 to 2025-07-01 5. Information about funding sources that supported the collection of the data: This work was supported by the BMBF-funded de.NBI Cloud within the German Network for Bioinformatics Infrastructure (de.NBI) (031A532B, 031A533A, 031A533B, 031A534A, 031A535A, 031A537A, 031A537B, 031A537C, 031A537D, 031A538A). 6. Language of the dataset: English 7. Geographic location of data collection: Germany B DATA & FILE OVERVIEW 1. File List: VB03.genome.fasta Victoria cruziana genome sequence at pseudochromosome level as FASTA file VB03.genome.pore-c_contact_map.png contact map generated by C-Phasing VB03.v001.anno.txt functional annotation results as TAB-separated file VB03.v001.cds.fasta predicted coding sequences as FASTA file VB03.v001.gff structural annotation results in GFF3 format VB03.v001.pep.fasta predicted polypeptide sequences as FASTA file md5sums.txt contains MD5 hashes of the other files 2. Are there multiple versions of the dataset? no 2.1 If yes, name of file(s) that was updated: i. Why was the file updated? ii. When was the file updated? 3. Relationship between files: md5sums.txt contains MD5 hashes and file name pairs of the other files 4. Additional related data collected that was not included in the current data package: n/a C SHARING/ACCESS INFORMATION 1. Was data derived from another source? Yes, functional annotation was derived from TAIR (https://www.arabidopsis.org/) 2. Licenses/restrictions placed on the data: CC BY 4.0 3. Links to publications that cite or use the data: Nowak M. S.*, Harder B.*, Meckoni S. N.*, Friedhoff R., Wolff K., Pucker B. (2024). Genome sequence and RNA-seq analysis reveal genetic basis of flower coloration in the giant water lily Victoria cruziana. bioRxiv 2024.06.15.599162; doi: 10.1101/2024.06.15.599162. 4. Links to other publicly accessible locations of the data: n/a 5. Links/relationships to ancillary datasets: n/a D METHODOLOGICAL INFORMATION 1. Description of methods used for collection/generation of data: The genome of a Victoria cruziana plant was sequenced with nanopore long reads (R9.4.1 and R10.4.1 flow cells, MinION Mk1B) based on high molecular weight DNA extracted with a previously established protocol (https://dx.doi.org/10.17504/protocols.io.bcvyiw7w). Additionally, Pore-C was conducted (https://dx.doi.org/10.17504/protocols.io.rm7vz9mmrgx1/v1). 2. Methods for processing the data: The genome sequence was assembled with Verkko2, scaffolding was conducted with C-Phasing and the gene models were predicted by BRAKER3 and GeMoMa, using additional hints derived from RNA-seq data of mainly flower samples. The functional annotation was predicted based on sequence similarity to well characterized Arabidopsis thaliana sequences. Please see our publication for additional details. 3. Instrument- and/or software-specific information needed to interpret the data: Genome, coding, and polypeptide sequences are provided in FASTA file format. The structural annotation is provided in the GFF3 format. The functional annotation is provided as a TAB-separated text file. All these files are accessible with basic text editors. Additionally, the Pore-C contact map is provided as a PNG file. 4. People involved in sample collection, processing, analysis and/or submission: all authors 5. Describe any quality-assurance procedures performed on the data: n/a 6. Standards and calibration information: FASTA, GFF3 7. Environmental/experimental conditions: The plant was grown at 28°C water temperature in a plant cultivation room under long day conditions (16h light per day) at room temperature.