README FILE This file was generated on May 8, 2026 by Boas Pucker. A. GENERAL INFORMATION Title of the dataset: Urtica dioica genome sequence and annotation Brief description of the research project and its aims: Here we present the genome sequence and annotation of Urtica dioica (commonly known as stinging nettle). Using ONT long-read sequencing data, a highly continuous genome sequence was generated. This data publication contains the genome sequence with corresponding structural and functional annotation. Author Information A. Investigator Contact Information Name: Katharina Wolff Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschallee 1, 53115 Bonn, Germany. Email: kwolff@uni-bonn.de B. Investigator Contact Information Name: Julie Anne Vieira Salgado de Oliveira Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschallee 1, 53115 Bonn, Germany. Email: jvieiras@uni-bonn.de C. Project Supervisor (Principal Investigator) Name: Prof. Dr. Boas Pucker Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschallee 1, 53115 Bonn, Germany. Email: pucker@uni-bonn.de D. In case of questions related to this dataset, please contact: Name: Prof. Dr. Boas Pucker Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschallee 1, 53115 Bonn, Germany. Email: pucker@uni-bonn.de Date of data collection: 2023-2026 Acknowledgments: This work was supported by the de.NBI Cloud within the German Network for Bioinformatics Infrastructure (de.NBI) and ELIXIR-DE (Forschungszentrum Jülich and W-de.NBI-001, W-de.NBI-004, W-de.NBI-008, W-de.NBI-010, W-de.NBI-013, W-de.NBI-014, W-de.NBI-016, W-de.NBI-022). We are grateful for the excellent support provided by the team of the University of Bonn Botanic Gardens. Language of the dataset: English B. DATA & FILE OVERVIEW Urtica_dioica.v1.genome.fasta.gz Genome sequence of Urtica dioica. Assembly was performed with Hifiasm. Urtica_dioica.v1.genome.gff.gz Structural annotation of the Urtica dioica genome sequence generated by GeMoMa v1.9. Urtica_dioica.v1.cds.fasta.gz Protein encoding sequences of Urtica dioica derived from the structural annotation of the genome sequence. Urtica_dioica.v1.pep.fasta.gz Polypeptide sequences of Urtica dioica inferred from the coding sequences. Urtica_dioica.v1.anno.txt.gz Functional annotation predicted for the polypeptide sequences of Urtica dioica based on information available about Arabidopsis thaliana sequences. C. SHARING/ACCESS INFORMATION Data derived from another source: RNA-seq reads for the generation of hints for the gene prediction have been deposited at the Sequence Read Archive (PRJEB63450). Licenses: CC BY 4.0 Links to related publications: There is an earlier version of the Urtica dioica genome sequence: https://leopard.tu-braunschweig.de/receive/dbbs_mods_00078250. Other publicly accessible locations: n/a D. METHODOLOGICAL INFORMATION People involved in data collection, processing, analysis and/or submission: all authors Collection of data: The genome of a Urtica dioica plant (XX-0-DE-0-BONN-49798) was sequenced with nanopore long reads (R9.4.1 flow cells on a MinION and R10.4.1 flow cells on a PromethION) based on high-molecular weight DNA extracted with a previously established protocol (https://dx.doi.org/10.17504/protocols.io.bcvyiw7w). The final assembly was generated based on the R10 data. Processing of data: Basecalling was done with dorado (ONT): MinION runs with dorado v1.0.2 and PromethION runs with dorado v1.2.0. Genome sequence assembly was done with Hifiasm v0.25.0-r726. The structural annotation was generated with GeMoMa v1.9 based on RNA-seq hints derived from RNA-seq read mapping with STAR v2.7.11b. The functional annotation was generated with construct_anno.py based on information available for well-characterized Arabidopsis thaliana sequences with high sequence similarity (homology). Data structure: Genome, coding, and polypeptide sequences are provided in FASTA file format. The structural annotation is provided in the GFF3 format. The functional annotation is provided as a TAB-separated text file. The functional annotation was predicted based on sequence similarity to well-characterized Arabidopsis thaliana sequences. The genome sequence was assembled with Hifiasm-0.25.0-r726. The gene models were predicted by GeMoMa v1.9. All files are accessible with basic text editors. File formats: TXT, FASTA, GFF