README FILE This file was generated on October 16th, 2025 by Boas Pucker. A. GENERAL INFORMATION Title of the dataset: Digitalis purpurea genome sequence and annotation V2 Brief description of the research project and its aims: Here we present a high-quality genome sequence and annotation of Digitalis purpurea (commonly known as foxglove), a plant valued both for its ornamental flowers and for producing the cardiac drug digoxin. The study aims to uncover the genetic basis of flower color variation in this species. Using ONT long-read sequencing data, we assembled the genome sequence of a magenta-flowering D. purpurea individual. We performed extensive gene prediction and functional annotation of protein-coding genes. Genes of the flavonoid biosynthesis were compared between magenta and white flowering individuals. Expression patterns were compared between different plant organs and between plants with different flower colors. Through these analyses, we identified a large insertion in the anthocyanidin synthase (ANS) that explains the differences between magenta and white flowers. This data publication contains the genome sequence with the corresponding annotation V2. Author Information A. Investigator Contact Information Name: Jakob Maximilian Horz Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: jahorz@uni-bonn.de Name: Katharina Wolff Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: k.wolff@uni-bonn.de Name: Ronja Friedhoff Institution: Plant Biotechnology and Bioinformatics, TU Braunschweig. Address: Mendelssohnstraße 4, 38106 Braunschweig, Germany. Email: r.friedhoff@tu-braunschweig.de B. Project Supervisor (Principal Investigator) Name: Prof. Dr. Boas Pucker Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: pucker@uni-bonn.de C. In case of questions related to this dataset, please contact: Name: Prof. Dr. Boas Pucker Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address:Kirschalle 1, 53115 Bonn, Germany. Email: pucker@uni-bonn.de Date of data collection: 2023 – 2025-10-12 Acknowledgments: This work was supported by the de.NBI Cloud within the German Network for Bioinformatics Infrastructure (de.NBI) and ELIXIR-DE (Forschungszentrum Jülich and W-de.NBI-001, W-de.NBI-004, W-de.NBI-008, W-de.NBI-010, W-de.NBI-013, W-de.NBI-014, W-de.NBI-016, W-de.NBI-022). Language of the dataset: English B. DATA & FILE OVERVIEW DR1_v1.annoV2.genome.fasta.gz Genome sequence of Digitalis purpurea. Assembly was performed with NextDenovo2. DNA originated from a magenta flowering plant. DR1_v1.annoV2.anno.gff.gz Structural annotation of the Digitalis purpurea genome sequence generated by GeMoMa v1.9. DR1_v1.annoV2.cds.fasta.gz Protein encoding sequences of Digitalis purpurea derived from the structural annotation of the genome sequence. DR1_v1.annoV2.pep.fasta.gz Polypeptide sequences of Digitalis purpurea inferred from the coding sequences. DR1_v1.annoV2.anno.txt.gz Functional annotation predicted for the polypeptide sequences of Digitalis purpurea based on information available about Arabidopsis thaliana sequences. DR1_v1.annoV2.tpms.txt.gz Transcript abundances (Transcripts Per Million, TPMs) calculated for the Digitalis purpurea genes. DR1_v1.annoV2.TEanno.gff3.gz Annotation of transposable elements for the Digitalis purpurea genome sequence generated with EDTA v2.2.2. C. SHARING/ACCESS INFORMATION Data derived from another source: RNA-seq reads for the generation of hints for the gene prediction have been retrieved from the Sequence Read Archive. Please see our corresponding preprint for additional details (https://doi.org/10.1101/2024.02.14.580303). Licenses: CC BY 4.0 Links to related publications: https://doi.org/10.1101/2024.02.14.580303 https://doi.org/10.24355/dbbs.084-202311130950-0 Other publicly accessible locations: n/a D. METHODOLOGICAL INFORMATION People involved in data collection, processing, analysis and/or submission: all authors Collection of data: The genome of a magenta flowering Digitalis purpurea plant was sequenced with nanopore long reads (R9.4.1 flow cells, MinION) based on high molecular weight DNA extracted with a previously established protocol (https://dx.doi.org/10.17504/protocols.io.bcvyiw7w). Processing of data: The genome sequence was assembled with NextDenovo2 and the gene models were predicted by GeMoMa v1.9. The functional annotation was predicted based on sequence similarity to well characterized Arabidopsis thaliana sequences. De novo annotation of transposable elements was performed with EDTA v2.2.2. Data structure: Genome, coding, and polypeptide sequences are provided in FASTA file format. The structural annotation and transposable elements annotation is provided in the GFF3 format. The functional annotation is provided as a TAB-separated text file. Gene expression data are provided as TAB-separated count table with samples in the first row and gene IDs in the first column. All files are accessible with basic text editors. File formats: TXT, FASTA, GFF