README FILE This file was generated on January 17th, 2026 by Julie Anne V.S. de Oliveira. A. GENERAL INFORMATION Title of the dataset: Purple and White Aquilegia vulgaris genome sequence and annotation Brief description of the research project and its aims: Here we present high-quality genome sequences and annotations of Aquilegia vulgaris (commonly known as columbine), a widespread ornamental plant renowned for its extensive flower color variation. The study uncovers the flavonoid biosynthesis gene repertoire underlying anthocyanin pigmentation in this species. Using ONT long-read sequencing data, highly continuous genome assemblies were generated for both a purple-flowering and a white-flowering A. vulgaris. This data publication contains the genome sequences with corresponding annotations. Author Information A. Investigator Contact Information Name: Julie Anne Vieira Salgado de Oliveira Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: jvieiras@uni-bonn.de Name: Ronja Friedhoff Institution: Plant Biotechnology and Bioinformatics, TU Braunschweig. Address: Mendelssohnstraße 4, 38106 Braunschweig, Germany. Email: r.friedhoff@tu-braunschweig.de Name: Katharina Wolff Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: k.wolff@uni-bonn.de B. Project Supervisor (Principal Investigator) Name: Prof. Dr. Boas Pucker Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address: Kirschalle 1, 53115 Bonn, Germany. Email: pucker@uni-bonn.de C. In case of questions related to this dataset, please contact: Name: Prof. Dr. Boas Pucker Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn. Address:Kirschalle 1, 53115 Bonn, Germany. Email: pucker@uni-bonn.de Date of data collection: 2023 – 2026-01-16 Acknowledgments: This work was supported by the de.NBI Cloud within the German Network for Bioinformatics Infrastructure (de.NBI) and ELIXIR-DE (Forschungszentrum Jülich and W-de.NBI-001, W-de.NBI-004, W-de.NBI-008, W-de.NBI-010, W-de.NBI-013, W-de.NBI-014, W-de.NBI-016, W-de.NBI-022). We are grateful for the excellent support provided by the team of the University of Bonn Botanic Gardens. Language of the dataset: English B. DATA & FILE OVERVIEW Aquilegia_vulgaris_purple.genomic.fasta.gz Genome sequence of Aquilegia vulgaris. Assembly was performed with NextDenovo2. DNA originated from a purple flowering plant. Aquilegia_vulgaris_purple.anno.gff.gz Structural annotation of the purple Aquilegia vulgaris genome sequence generated by GeMoMa v1.9. Aquilegia_vulgaris_purple.cds.fasta.gz Protein encoding sequences of purple Aquilegia vulgaris derived from the structural annotation of the genome sequence. Aquilegia_vulgaris_purple.pep.fasta.gz Polypeptide sequences of purple Aquilegia vulgaris inferred from the coding sequences. Aquilegia_vulgaris_purple.anno.txt.gz Functional annotation predicted for the polypeptide sequences of purple Aquilegia vulgaris based on information available about Arabidopsis thaliana sequences. Aquilegia_vulgaris_white.genomic.fasta.gz Genome sequence of Aquilegia vulgaris. Assembly was performed with Hifiasm. DNA originated from a white flowering plant. Aquilegia_vulgaris_white.anno.gff.gz Structural annotation of the white Aquilegia vulgaris genome sequence generated by GeMoMa v1.9. Aquilegia_vulgaris_white.cds.fasta.gz Protein encoding sequences of white Aquilegia vulgaris derived from the structural annotation of the genome sequence. Aquilegia_vulgaris_white.pep.fasta.gz Polypeptide sequences of white Aquilegia vulgaris inferred from the coding sequences. Aquilegia_vulgaris_white.anno.txt.gz Functional annotation predicted for the polypeptide sequences of white Aquilegia vulgaris based on information available about Arabidopsis thaliana sequences. C. SHARING/ACCESS INFORMATION Data derived from another source: RNA-seq reads for the generation of hints for the gene prediction have been retrieved from the Sequence Read Archive. Please see our corresponding preprint for additional details (https://doi.org/10.1101/2024.12.16.628782). Licenses: CC BY 4.0 Links to related publications: https://doi.org/10.1101/2024.12.16.628782 A first version of the dataset can be found under the following link: https://doi.org/10.24355/dbbs.084-202410280909-0 Other publicly accessible locations: n/a D. METHODOLOGICAL INFORMATION People involved in data collection, processing, analysis and/or submission: all authors Collection of data: The genome of a purple flowering Aquilegia vulgaris (TU Braunschweig-49638/1) plant was sequenced with nanopore long reads (R9.4.1 flow cells, MinION); the genome of a white flowering Aquilegia vulgaris plant was sequenced with nanopore long reads (R10.4.1 flow cells, MinION) based on high molecular weight DNA extracted with a previously established protocol (https://dx.doi.org/10.17504/protocols.io.bcvyiw7w). Processing of data: The genome sequence was assembled with NextDenovo2 for the purple one, and Hifiasm-0.25.0-r726 for the white one. The gene models were predicted by GeMoMa v1.9 for both. The functional annotation was predicted based on sequence similarity to well characterized Arabidopsis thaliana sequences. Data structure: Genome, coding, and polypeptide sequences are provided in FASTA file format. The structural annotation is provided in the GFF3 format. The functional annotation is provided as a TAB-separated text file. All files are accessible with basic text editors. File formats: TXT, FASTA, GFF