README FILE This file was generated on July 31th, 2026 by Julie Anne V.S. de Oliveira. A. GENERAL INFORMATION Title of the dataset: Valeriana officinalis genome sequence and annotation V1 Brief description of the research project and its aims: Here we present the first genome sequence and annotation of Valeriana officinalis (commonly known as valerian), a widespread perennial plant widely used as a herbal remedy. Using ONT long-read sequencing data, a highly continuous genome assembly was generated for V. officinalis. This study uncovers the flavonoid biosynthesis gene repertoire and numerous terpene synthase candidate genes in this species. This data publication contains the genome sequences with corresponding annotations. Author Information A. Investigator Contact Information Name: Julie Anne Vieira Salgado de Oliveira Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: jvieiras@uni-bonn.de Name: Mariana Alejandra Baez Institution: Plant Breeding Department, Institute of Crop Science and Resource Conservation - INRES, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: mbaez@uni-bonn.de B. Project Supervisor (Principal Investigator) Name: Prof. Dr. Boas Pucker Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: pucker@uni-bonn.de C. In case of questions related to this dataset, please contact: Name: Prof. Dr. Boas Pucker Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address:Kirschalle 1, 53115 Bonn, Germany. Email: pucker@uni-bonn.de Date of data collection: 2026-07-31 Acknowledgments: This work was supported by the de.NBI Cloud within the German Network for Bioinformatics Infrastructure (de.NBI) and ELIXIR-DE (Forschungszentrum Jülich and W-de.NBI-001, W-de.NBI-004, W-de.NBI-008, W-de.NBI-010, W-de.NBI-013, W-de.NBI-014, W-de.NBI-016, W-de.NBI-022). We are grateful for the excellent support provided by the team of the University of Bonn Botanic Gardens. Language of the dataset: English B. DATA & FILE OVERVIEW Valeriana_officinalis.genomic.fasta.gz Genome sequence of Valeriana officinalis in a gzip-compressed FASTA file. The assembly was performed with Shasta and Hifiasm-0.25.0-r726, respectively. A whole-genome sequence alignment between the two assemblies was performed with nucmer (MUMmer4) and subjected to quickmerge v0.3 to merge the two assemblies into one final sequence. Valeriana_officinalis.anno.gff.gz Structural annotation of the Valeriana officinalis genome sequence generated by GeMoMa v1.9. Valeriana_officinalis.cds.fasta.gz Protein encoding sequences of Valeriana officinalis derived from the structural annotation of the genome sequence. Valeriana_officinalis.pep.fasta.gz Polypeptide sequences of Valeriana officinalis inferred from the coding sequences. Valeriana_officinalis.anno.txt.gz Functional annotation predicted for the polypeptide sequences of Valeriana officinalis based on gene function information available for Arabidopsis thaliana sequences. C. SHARING/ACCESS INFORMATION Data derived from another source: RNA-seq reads for the generation of hints for the gene prediction have been retrieved from the Sequence Read Archive. Please see our corresponding preprint for additional details (https://doi.org/10.64898/2026.08.14.744958). Licenses: CC BY 4.0 Other publicly accessible locations: n/a D. METHODOLOGICAL INFORMATION People involved in data collection, processing, analysis and/or submission: all authors Collection of data: The genome of a Valeriana officinalis plant originating from the seeds XX-0-BRAUN-7422401 (TU Braunschweig Botanical Garden) was sequenced with nanopore long reads (R10.4.1 flow cells, PromethION) based on high molecular weight DNA extracted with a previously established protocol (https://dx.doi.org/10.17504/protocols.io.bcvyiw7w). Processing of data: The genome sequence was assembled with Shasta and Hifiasm-0.25.0-r726, respectively. A whole-genome sequence alignment between the two assemblies was performed with nucmer (MUMmer4) and subjected to quickmerge v0.3 to merge the two assemblies into one final sequence. The gene models were predicted by GeMoMa v1.9. The functional annotation was predicted based on sequence similarity to well characterized Arabidopsis thaliana sequences. Data structure: Genome sequence, coding sequences, and polypeptide sequences are provided in FASTA file format (gzip-compressed). The structural annotation is provided in the GFF format. The functional annotation is provided as a TAB-separated text file. All files are accessible with basic text editors. File formats: TXT, FASTA, GFF