README FILE This file was generated on April 28th, 2026 by Katharina Wolff. A. GENERAL INFORMATION Title of the dataset: Rubus armeniacus genome sequence and annotation Brief description of the research project and its aims: Here, we present the genome sequence and annotation of Rubus armeniacus. Oxford Nanopore Technologies (ONT) long-read sequencing data were assembled to generate the genome sequence. Protein-coding genes were predicted and functionally annotated. Author Information A. Investigator Contact Information Name: Katharina Wolff Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: kwolff@uni-bonn.de B. Project Supervisor (Principal Investigator) Name: Prof. Dr. Boas Pucker Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: pucker@uni-bonn.de C. In case of questions related to this dataset, please contact: Name: Prof. Dr. Boas Pucker Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address:Kirschalle 1, 53115 Bonn, Germany. Email: pucker@uni-bonn.de Date of data collection: 2024 – 2026 Acknowledgments: This work was supported by the de.NBI Cloud within the German Network for Bioinformatics Infrastructure (de.NBI) and ELIXIR-DE (Forschungszentrum Jülich and W-de.NBI-001, W-de.NBI-004, W-de.NBI-008, W-de.NBI-010, W-de.NBI-013, W-de.NBI-014, W-de.NBI-016, W-de.NBI-022). Language of the dataset: English B. DATA & FILE OVERVIEW Rubus_armeniacus.genome.fasta Genome sequence of Rubus armeniacus. The assembly was performed with Shasta. DNA was obtained from leaves of multiple clones of a Rubus armeniacus plant (DE-0-BONN-49806). Rubus_armeniacus.hapA.fasta Pseudohaplophase A of Rubus armeniacus based on similarity to Rubus caesius, extracted from Rubus_armeniacus.genome.fasta Rubus_armeniacus.hapB.fasta Pseudohaplophase B of Rubus armeniacus based on similarity to Rubus caesius, extracted from Rubus_armeniacus.genome.fasta Rubus_armeniacus.hapC.fasta Pseudohaplophase C of Rubus armeniacus based on similarity to Rubus caesius, extracted from Rubus_armeniacus.genome.fasta Rubus_armeniacus.hapD.fasta Pseudohaplophase D of Rubus armeniacus based on similarity to Rubus caesius, extracted from Rubus_armeniacus.genome.fasta Rubus_armeniacus.anno.gff Structural annotation of the Rubus armeniacus genome sequence generated by GeMoMa v1.9. and BRAKER3. Rubus_armeniacus.cds.fasta Protein encoding sequences of Rubus armeniacus derived from the structural annotation of the genome sequence. Rubus_armeniacus.pep.fasta Polypeptide sequences of Rubus armeniacus inferred from the coding sequences. Rubus_armeniacus.anno.txt Functional annotation predicted for the polypeptide sequences of Rubus armeniacus based on information available about Arabidopsis thaliana sequences. Rubus_armeniacus.fun.anno.txt Functional annotation predicted for the polypeptide sequences of Rubus armeniacus based on funannotate with the InterProScan5 and Pfam database. Rubus_armeniacus.tpms.txt Transcript abundances (Transcripts Per Million, TPMs) calculated for the Rubus armeniacus genes. C. SHARING/ACCESS INFORMATION Data derived from another source: n/a Licenses: CC BY 4.0 Links to related publications: n/a Other publicly accessible locations: n/a D. METHODOLOGICAL INFORMATION People involved in data collection, processing, analysis and/or submission: all authors Collection of data: Leaves of multiple Rubus armeniacus clones were used to extract high-molecular-weight DNA via an established protocol (https://dx.doi.org/10.17504/protocols.io.bcvyiw7w), followed by nanopore long-read sequencing (MinION Mk1B). Processing of data: The genome sequence was assembled with Shasta and the gene models were predicted by GeMoMa v1.9 and BRAKER3. The functional annotation was predicted based on sequence similarity to well characterized Arabidopsis thaliana sequences, as well as the tool funannotate utilizing InterProScan5 and Pfam as a database. Data structure: Genome, coding, and polypeptide sequences are provided in FASTA file format. The structural annotation is provided in the GFF3 format. The functional annotation is provided as a TAB-separated text file. Gene expression data (Rubus_armeniacus.tpms.tsv) are provided as TAB-separated count table with samples in the first row and gene IDs in the first column. Samples B1 to B6 are based on RNAseq-data from dark berries, samples R1 to R6 are based on RNA-seq data from red berries, H1 to H6 are based on RNAseq-data from half-red berries, G1 to G6 are based on RNAseq-data from green berries, and L1 to L3 are based on RNAseq-data from green leaves. All files are accessible with basic text editors. File formats: TXT, FASTA, GFF, TSV