This file was generated on 2026-06-17 by Samuel Nestor Meckoni


# A GENERAL INFORMATION

1. Title of the dataset:
    ```
	Utricularia gibba genome sequence and annotation
    ```

2. Brief description of the research project and its aims:
	```
	The genome of an Utricularia gibba plant was sequenced with nanopore long reads. The genome sequence was assembled with hifiasm and the gene models were predicted by Helixer, BRAKER3 and GeMoMa. The functional annotation was predicted based on sequence similarity to well characterized Arabidopsis thaliana sequences.
    ```

3. Author Information
    - A. Investigator Contact Information
        ```
		Name: Samuel Nestor Meckoni
		Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn.
		Address: Kirschalle 1, 53115 Bonn, Germany
		Email: meckoni@uni-bonn.de
        ```

	- B. Investigator Contact Information
        ```
		Name: Julie Anne Vieira Salgado de Oliveira
		Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn
		Address: Kirschalle 1, 53115 Bonn, Germany
		Email: jvieiras@uni-bonn.de
        ```
		
	- C. Project Supervisor (Principal Investigator) Contact Information
        ```
		Name: Prof. Dr. Boas Pucker
		Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn
		Address: Kirschalle 1, 53115 Bonn, Germany
		Email: pucker@uni-bonn.de
        ```

	- D. In case of questions related to this dataset, please contact:
        ```
		Name: Prof. Dr. Boas Pucker
		Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn
		Address: Kirschalle 1, 53115 Bonn, Germany
		Email: pucker@uni-bonn.de
        ```
        
4. Date of data collection:
    ```
	2025-07-01 to 2025-12-31
	```

5. Information about funding sources that supported the collection of the data:
    ```
	This work was supported by the de.NBI Cloud within the German Network for Bioinformatics Infrastructure (de.NBI) and ELIXIR-DE (Forschungszentrum Jülich and W-de.NBI-001, W-de.NBI-004, W-de.NBI-008, W-de.NBI-010, W-de.NBI-013, W-de.NBI-014, W-de.NBI-016, W-de.NBI-022). We are grateful for the excellent support provided by the team of the University of Bonn Botanic Gardens.
    ```

6. Language of the dataset:
	```
    English
    ```



# B DATA & FILE OVERVIEW

1. File List:
	```
	- daUtrGibb.v01.asm.fasta.gz
		- Utricularia gibba genome sequence as FASTA file

	- daUtrGibb.v01.v01.anno.txt.gz
		- functional annotation results as TAB-separated file

	- daUtrGibb.v01.v01.cds.fasta.gz
		- predicted coding sequences as FASTA file

	- daUtrGibb.v01.v01.gff.gz
		- structural annotation results in GFF3 format

	- daUtrGibb.v01.v01.longest.isoforms.cds.fasta.gz
		- predicted coding sequences as FASTA file, only for the longest isoform per predicted gene

	- daUtrGibb.v01.v01.longest.isoforms.gff.gz
		- structural annotation results in GFF3 format, only for the longest isoform per predicted gene

	- daUtrGibb.v01.v01.longest.isoforms.pep.fasta.gz
		- predicted polypeptide sequences as FASTA file, only for the longest isoform per predicted gene

	- daUtrGibb.v01.v01.pep.fasta.gz
		- predicted polypeptide sequences as FASTA file

	- md5sums.txt
		- contains MD5 hashes of the uncompressed files

	- md5sums_gzipped.txt
	    - contains MD5 hashes of the gzip compressed files
    ```

2. Are there multiple versions of the dataset?
    ```
	no
	```

3. Relationship between files:
	```
    md5sums_gzipped.txt contains MD5 hashes and file name pairs of the gzip compressed files
    md5sums.txt contains MD5 hashes and file name pairs of the uncompressed files
    ```

4. Additional related data collected that was not included in the current data package:
	```
    n/a
    ```



# C SHARING/ACCESS INFORMATION

1. Was data derived from another source?
	```
    yes
	Functional annotation was derived from TAIR (https://www.arabidopsis.org/)
	Structural annotation was supported by RNA-seq data
	```

2. Licenses/restrictions placed on the data:
    ```
    CC BY 4.0
    ```

3. Links to publications that cite or use the data:
	```
    n/a
    ```

4. Links to other publicly accessible locations of the data:
	```
    n/a
    ```

5. Links/relationships to ancillary datasets:
	```
    n/a
    ```



# D METHODOLOGICAL INFORMATION

1. Description of methods used for collection/generation of data: 
	```
    The genome of an Utricularia gibba plant was sequenced with nanopore long reads (R10.4.1 flow cells, MinION Mk1B) based on high molecular weight DNA extracted with a previously established protocol (https://dx.doi.org/10.17504/protocols.io.bcvyiw7w).
    ```

2. Methods for processing the data: 
	```
    The genome sequence was assembled with hifiasm (version: 0.25.0-r726) and the gene models were predicted by Helixer (version: 0.3.5), BRAKER3 (version: v3.0.8), and GeMoMa (version: 1.9, build hdfd78af_0, bioconda), using additional hints derived from RNA-seq data. The functional annotation was predicted based on sequence similarity to well characterized Arabidopsis thaliana sequences. Longest isoforms have been filtered with AGAT agat_sp_keep_longest_isoform.pl (version: 1.4.1).
    ```

3. Instrument- and/or software-specific information needed to interpret the data: 
	```
    Genome, coding, and polypeptide sequences are provided in FASTA file format. The structural annotation is provided in the GFF3 format. The functional annotation is provided as a TAB-separated text file. All these files are accessible with basic text editors.
    ```

4. People involved in sample collection, processing, analysis and/or submission:
	```
    all authors
    ```

5. Describe any quality-assurance procedures performed on the data:
	```
    n/a
    ```

6. Standards and calibration information:
	```
    FASTA, GFF3
    ```

7. Environmental/experimental conditions:
	```
    The plant cultures originate from Utricularia gibba (XX-0-DATH-519) and were grown in in vitro cultures with ca. 2400-2700 lx 16/8h day/night light (Philips TL5 39W/840 HO) and 22-25 °C in glass jars off 370 ml size (WECK-Sturzglas RR100) that were filled with ca. 75 ml 1/4 MS2 + 0.25 mg/l 6-benzylaminopurine (BAP).
    ```
