This file was generated on 2026-08-17 by Maria F. Marin Recinos


# A GENERAL INFORMATION

1. Title of the dataset:
    ```
	Theobroma cacao: Genome Sequences and Annotations
    ```

2. Brief description of the research project and its aims:
	```
	The genomes of two Theobroma cacao plants were sequenced with Nanopore long reads. The genome sequence was assembled with hifiasm and the gene models were predicted by GeMoMa. The functional annotation was predicted based on sequence similarity to well characterized Arabidopsis thaliana sequences.
    ```

3. Author Information
	- A. Investigator Contact Information
        ```
		Name: Maria F. Marin Recinos
		Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn.
		Address: Kirschalle 1, 53115 Bonn, Germany
		Email: marmarin@uni-bonn.de
        ```
        
	- B. Investigator Contact Information
        ```
		Name: Katharina Wolff
		Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn
		Address: Kirschalle 1, 53115 Bonn, Germany
		Email: kwolff@uni-bonn.de
        ```
        
	- C. Investigator Contact Information
        ```
		Name: Nancy Choudhary
		Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn
		Address: Kirschalle 1, 53115 Bonn, Germany
		Email: n.choudhary@uni-bonn.de
        ```
        
	- D. Investigator Contact Information
        ```
		Name: Julie Anne Vieira Salgado de Oliveira
		Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn
		Address: Kirschalle 1, 53115 Bonn, Germany
		Email: jvieiras@uni-bonn.de
        ```		
        
	- C. Project Supervisor (Principal Investigator) Contact Information
        ```
		Name: Prof. Dr. Boas Pucker
		Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn
		Address: Kirschalle 1, 53115 Bonn, Germany
		Email: pucker@uni-bonn.de
        ```

	- D. In case of questions related to this dataset, please contact:
        ```
		Name: Prof. Dr. Boas Pucker
		Institution: Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany - IZMB, University of Bonn
		Address: Kirschalle 1, 53115 Bonn, Germany
		Email: pucker@uni-bonn.de
        ```
        
4. Date of data collection:
    ```
	2024-05-01 to 2025-12-31
	```

5. Information about funding sources that supported the collection of the data:
    ```
	This work was supported by the de.NBI Cloud within the German Network for Bioinformatics Infrastructure (de.NBI) and ELIXIR-DE (Forschungszentrum Jülich and W-de.NBI-001, W-de.NBI-004, W-de.NBI-008, W-de.NBI-010, W-de.NBI-013, W-de.NBI-014, W-de.NBI-016, W-de.NBI-022). We thank all members of the research group Plant Biotechnology and Bioinformatics for discussion and support. We are grateful for the excellent support provided by the team of the University of Bonn Botanic Gardens.
    ```

6. Language of the dataset:
	```
    English
    ```



# B DATA & FILE OVERVIEW


1. File List:
	```
	- Tcacao_BRAU1.v01.asm.fasta.gz
		- Theobroma cacao BRAU1 genome assembly in compressed FASTA format.
		
	- Tcacao_BRAU1.v01.v01.gff
		- Structural gene annotation of the BRAU1 genome sequence in GFF3 format.

	- Tcacao_BRAU1.v01.v01.cds.fasta
		- Predicted coding sequences (CDS) of the BRAU1 genome sequence in FASTA format.

	- Tcacao_BRAU1.v01.v01.pep.fasta
		- Predicted polypeptide sequences of the BRAU1 genome sequence in FASTA format.
		
	- Tcacao_BRAU1_annot.v01.v01.txt
		- Functional annotation of predicted BRAU1 polypeptides in tab-separated text format.

	- Tcacao_BONN1.v01.asm.fasta.gz
		- Theobroma cacao BONN1 genome assembly in compressed FASTA format.

	- Tcacao_BONN1.v01.v01.gff
		- Structural gene annotation of the BONN1 genome sequence in GFF3 format.

	- Tcacao_BONN1.v01.v01.cds.fasta
		- Predicted coding sequences (CDS) of the BONN1 genome sequence in FASTA format.

	- Tcacao_BONN1.v01.v01.pep.fasta
		- Predicted polypeptide sequences of the BONN1 genome sequence in FASTA format.
		
	- Tcacao_BONN1_annot.v01.v01.txt
		- Functional annotation of predicted BONN1 polypeptides in tab-separated text format.
    ```
    

2. Are there multiple versions of the dataset?
    ```
	no
	```

3. Additional related data collected that was not included in the current data package:
	```
    n/a
    ```



# C SHARING/ACCESS INFORMATION

1. Was data derived from another source?
	```
    yes
	Functional annotation was derived from TAIR (https://www.arabidopsis.org/)
	Structural annotation was supported by RNA-seq data (SRR23566372, SRR23566373, SRR23566377, SRR23584463, SRR23584464, SRR23584468)
	```

2. Licenses/restrictions placed on the data:
    ```
    CC BY 4.0
    ```

3. Links to publications that cite or use the data:
	```
    https://doi.org/10.1101/2024.11.23.624982
    ```

4. Links to other publicly accessible locations of the data:
	```
    n/a
    ```

5. Links/relationships to ancillary datasets:
	```
    Sequencing data have been deposited at the European Nucleotide Archive:
    https://www.ebi.ac.uk/ena/browser/view/PRJEB65101
    https://www.ebi.ac.uk/ena/browser/view/PRJEB97086
    ```



# D METHODOLOGICAL INFORMATION

1. Description of methods used for collection/generation of data: 
	```
    Theobroma cacao trees XX-0-BRAUN-7843067 and XX-0-BONN-19718-6-1990 were used for generation of the BRAU1 and BONN1 genome assemblies, respectively. The trees have been cultivated under greenhouse conditions for several decades at the Botanical Garden of TU Braunschweig and the Botanical Gardens of the University of Bonn. The geographic origin of both trees is unknown. Fresh and healthy leaves were harvested on the same day as DNA extraction to ensure high DNA quantity and quality. For the BONN1 sample, DNA was extracted using a previously described CTAB-based protocol (Siadjeu et al., 2020). For the BRAU1 sample, DNA extraction was performed using a modified CTAB-based protocol. Fresh leaves were first cleared of midribs and major veins, leaving predominantly foliar tissue to reduce the content of polyphenols and polysaccharides. Approximately 0.9 g of foliar tissue was ground to a fine powder using liquid nitrogen. Following DNA extraction, DNA quality and quantity were assessed using NanoDrop measurements, agarose gel electrophoresis, and Qubit fluorometric measurements. High-molecular weight DNA was subsequently prepared for long-read sequencing. For the BRAU1 and BONN1 genomes, sequencing data were generated using long-read sequencing technology. The resulting sequencing data were used for de novo genome assembly with hifiasm.
    ```

2. Methods for processing the data: 
	```
    The genome sequences of BRAU1 and BONN1 were assembled de novo using hifiasm version 0.25.0-r726. Structural gene annotation was performed using GeMoMa version 1.9, using RNA-seq-derived hints to support gene model prediction. The resulting gene models were used to generate predicted coding sequences and polypeptide sequences. A general functional annotation of predicted polypeptide sequences was performed using the customized Python script construct_anno.py (Pucker & Iorizzo, 2023), based on sequence similarity to well-characterized Arabidopsis thaliana proteins.
    ```

3. Instrument- and/or software-specific information needed to interpret the data: 
	```
    Genome assemblies, coding sequences, and predicted polypeptide sequences are provided in FASTA format. Structural gene annotations are provided in GFF3 format. Functional annotations are provided as tab-separated text files. The genome assemblies were generated using hifiasm version 0.25.0-r726. Structural gene annotation was performed using GeMoMa version 1.9. Functional annotation was performed using construct_anno.py (Pucker & Iorizzo, 2023). FASTA and GFF3 files can be accessed using standard sequence-analysis and genome-annotation software. Tab-separated annotation files can be accessed using standard spreadsheet or text-processing software.
    ```

4. People involved in sample collection, processing, analysis and/or submission:
	```
	All authors were involved in the generation, processing, analysis, and/or submission of the dataset.
    ```

5. Describe any quality-assurance procedures performed on the data:
	```
    n/a
    ```

6. Standards and calibration information:
	```
    FASTA, GFF3
    ```

7. Environmental/experimental conditions:
	```
    The Theobroma cacao trees XX-0-BRAUN-7843067 and XX-0-BONN-19718-6-1990 used in this study have been cultivated under greenhouse conditions at the Botanical Garden of TU Braunschweig and the Botanical Gardens of the University of Bonn, respectively, for several decades. The geographic origin of both trees is unknown. Fresh and healthy leaves were harvested from both trees on the same day as DNA extraction.
    ```
