README FILE This file was generated on January 26th, 2026 by Victoria Fischer. A. GENERAL INFORMATION Title of the dataset: Begonia manicata genome sequence and annotation Brief description of the research project and its aims: Here, we present the genome sequence and annotation of Begonia manicata. Oxford Nanopore Technologies (ONT) long-read sequencing data were assembled to generate the genome sequence. Protein-encoding genes were predicted and functionally annotated. Author Information A. Investigator Contact Information Name: Victoria Fischer Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: s09vfisc@uni-bonn.de Name: Chiara Marie Dassow Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: s67cdass@uni-bonn.de B. Project Supervisor (Principal Investigator) Name: Prof. Dr. Boas Pucker Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: pucker@uni-bonn.de C. In case of questions related to this dataset, please contact: Name: Prof. Dr. Boas Pucker Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address:Kirschalle 1, 53115 Bonn, Germany. Email: pucker@uni-bonn.de Date of data collection: 2025 – 2026-01-21 Acknowledgments: This work was supported by the de.NBI Cloud within the German Network for Bioinformatics Infrastructure (de.NBI) and ELIXIR-DE (Forschungszentrum Jülich and W-de.NBI-001, W-de.NBI-004, W-de.NBI-008, W-de.NBI-010, W-de.NBI-013, W-de.NBI-014, W-de.NBI-016, W-de.NBI-022). Language of the dataset: English B. DATA & FILE OVERVIEW Begonia_manicata.genome.fasta Genome sequence of Begonia manicata. The assembly was performed with Hifiasm. DNA was obtained from leaves of multiple clones of a Begonia manicata plant (XX-0-BRAUN-7104585). Begonia_manicata.anno.gff Structural annotation of the Begonia manicata genome sequence generated by GeMoMa v1.9. Begonia_manicata.cds.fasta Protein encoding sequences of Begonia manicata derived from the structural annotation of the genome sequence. Begonia_manicata.pep.fasta Polypeptide sequences of Begonia manicata inferred from the coding sequences. Begonia_manicata.anno.txt Functional annotation predicted for the polypeptide sequences of Begonia manicata based on information available about Arabidopsis thaliana sequences. Begonia_manicata.tpms.txt Transcript abundances (Transcripts Per Million, TPMs) calculated for the Begonia manicata genes. C. SHARING/ACCESS INFORMATION Data derived from another source: n/a Licenses: CC BY 4.0 Links to related publications: n/a Other publicly accessible locations: n/a D. METHODOLOGICAL INFORMATION People involved in data collection, processing, analysis and/or submission: all authors Collection of data: Leaves of multiple Begonia manicata clones were used to extract high-molecular-weight DNA via an established protocol (https://dx.doi.org/10.17504/protocols.io.bcvyiw7w), followed by nanopore long-read sequencing (MinION MK1B). Red structures emerging from leaves and stem were only harvested for RNA-seq. Processing of data: The genome sequence was assembled with Hifiasm and the gene models were predicted by GeMoMa v1.9. The functional annotation was predicted based on sequence similarity to well characterized Arabidopsis thaliana sequences. Data structure: Genome, coding, and polypeptide sequences are provided in FASTA file format. The structural annotation is provided in the GFF3 format. The functional annotation is provided as a TAB-separated text file. Gene expression data (Begonia_manicata.tpms.txt) are provided as TAB-separated count table with samples in the first row and gene IDs in the first column. While samples 01 to 05 are based on RNAseq-data from leave tissue, samples 06 to 10 are based on RNA-seq data from red tissue on stems. All files are accessible with basic text editors. File formats: TXT, FASTA, GFF