Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.
2,095 rows · 907 KB · fasta, tsv
A virus genome database representing 21,015 near-complete virus genomes collected from untargeted ultra-deep RNA/DNA combined sequencing of wastewater. Sequence data was provided by the CASPER consortium and raw data may be found on NCBI SRA under bioprojects PRJNA1247874 and PRJNA1198001. Data underwent read trimming, rRNA and human read removal, de novo assembly, and selection of high-quality viral contigs. Contigs were clustered at 95% identity and 85% query coverage to dereplicate. Chimera-checking required at least two independent assemblies of the same viral genome or presence of the genome in another reference database. Annotation made use of RdRpCATCH, geNomad, checkV, BLASTN against NCBI core-nt, and RNAVirHost. The RdRp fasta files contain representative RdRp sequences identified through homology to major RdRp reference databases and clustered at 90% sequence identity over 75% sequence coverage. Included sequences contain all three conserved RdRp motifs (A, B, and C) arranged in either the canonical ABC configuration or the permuted CAB configuration.
Aligned Fasta files from Fianco & Melo (in press) Evolution of the false-leaf katydids (Orthoptera: Tettigoniidae: Phaneropterinae): molecular phylogeny and divergence times
Hubbard, Alfred · Solares, Edwin · Hemming-Schroeder, Elizabeth
88 rows · 18 KB · fasta
These are the files needed to run the Broad Institute's malaria amplicon pipeline for the PvGAP Plasmodium vivax panel, described in detail here . They consist of FASTA files containing the forward and reverse primers and another FASTA file containing reference sequences for each target, derived from the PvP01 reference genome.
This dataset contains the FASTA sequences of PCR amplicons generated using the multiplex PCR primers described in Khokhar et al ., "A low resource requirement molecular diagnostic and surveillance tool for Shigella in the era of vaccination" (submitted to Communications Medicine). The 15 amplicon sequences correspond to the Shigella serotyping targets ( gtr I, gtr II, gtr X, oac -1b, oac -3a), species identification targets ( ipaH , Shigella sonnei DNA methylase), antimicrobial resistance genes ( bla CTX-M variants 1, 3, 14, 15, 27, 55; mph( A)), and the human RNaseP internal control. These sequences were used to construct a custom ARIBA database for in silico validation of mPCR amplicon specificity against whole genome sequencing data.
This dataset contains the input sequences and the complete bioinformatic output files supporting the analysis of the Cannabis sativa Monoecy1 sex-determination locus reported in the associated manuscript. All analyses were performed on the Cannabis sativa cv. 'Pink Pepper' genome assembly GCA_029168945.1 (ASM2916894v1), chromosome X (RefSeq accession NC_083610.1), using the Dardel high-performance computing system at the PDC Center for High Performance Computing, KTH Royal Institute of Technology, Stockholm. The files document four lines of evidence used to evaluate the mechanism by which the long non-coding RNA lncREM16 (LOC133032448; transcript XR_009685085.1) silences the B3-domain transcription factor gene CsREM16 (LOC115699937; transcripts XM_030627485.2 and XM_030627490.2) within the Monoecy1 locus. Contents: Input sequences (FASTA): the 80 kb Monoecy1 genomic locus (Monoecy1_locus.fa; NC_083610.1:81,080,000-81,160,000), the two CsREM16 mRNA isoforms (CsREM16_X1.fa, CsREM16_X2.fa, CsREM16_both.fa), the lncREM16 transcript (lncREM16.fa), and the extracted 73 bp shared region from both the lncREM16 transcript and the genomic locus (shared_73bp.fa, shared_genomic.fa, both_shared.fa). Transposable element annotation (RepeatMasker v4.1.5, RMBLAST v2.14.0+, Dfam 3.7, viridiplantae): annotation tables, summary statistics, masked sequences, category files, and GFF output for both the 80 kb locus (Monoecy1_locus.fa.out, .tbl, .cat, .masked, .out.gff) and the lncREM16 transcript (lncREM16.fa.out, .tbl, .cat, .masked, .out.gff, .out.html). These files report the 33% TE content of the locus and the 37% TE content of the lncREM16 transcript, and identify the unclassified element DR2330744 at the CsREM16 exon 3 / intron 2 boundary. Sequence complementarity tests (NCBI blastn): antisense BLAST of lncREM16 against the CsREM16 mRNAs (lnc_vs_REM16_antisense.txt), against the genomic locus (lnc_vs_genomic_antisense.txt), the corresponding sense-strand orientation check (lnc_vs_genomic_sense.txt), the lncREM16 self-BLAST for inverted-repeat / hairpin detection (lnc_selfblast.txt), the alignment of the 73 bp shared region (shared_vs_lnc.txt, shared_vs_rem16.txt), and the BLAST of the 73 bp query against the NCBI nt database (blast_results.xml, blast_rid.txt). All antisense complementarity tests returned zero hits. Small RNA mapping (Bowtie v1): size-filtered small RNA reads (18-30 nt) derived from NCBI SRA run SRR25938888 (gerola_18_30.fastq.gz), and the alignment outputs in antisense and sense orientation against the CsREM16 mRNAs and the genomic locus (antisense_CsREM16.sam, sense_CsREM16.sam, antisense_monoecy1.sam, sense_monoecy1.sam) with their mapping statistics (antisense_stats.txt, sense_stats.txt, antisense_monoecy1_stats.txt, sense_monoecy1_stats.txt, overall_mapping_stats.txt, and the relaxed three-mismatch run antisense_v3.sam, antisense_v3_stats.txt). Gene records: NCBI gene-information records documenting the assembly, coordinates, and exon structure of CsREM16 and lncREM16 (rem16_gene_info.txt, lnc_gene_info.txt). Software versions and parameters are documented in the associated manuscript. The preprint / published article DOI will be added to this record under related identifiers upon availability.
Dataset: Eight complete chloroplast genomes of Oreocharis guileana 1. Sample list: SZ2_2, SZ2_3, SZ2_4, HK3_1, HK3_6, HK3_12, HK3_18, HK2_1 2. Data content: Only assembled FASTA sequences (no gene annotation files) 3. Usage: These sequences are used for phylogenetic tree reconstruction
This dataset contains the curated FASTA file of protein sequences used to construct a custom DIAMOND database for homology-based screening of dehalogenation-, putative defluorination-, and fluoride metabolism-related genes in bacterial genomes. The dataset includes sequence identifiers and amino acid sequences retrieved from UniProt and manually curated prior to DIAMOND database construction.
Requena Rolanía, Jose María · Greif, Gonzalo · ROBELLO, CARLOS
1 files · 8.0 MB · fasta
This dataset contains the genome sequence for Trypanosoma cruzi (strain Dm28c). This genome sequence was de novo assembled using PacBio Hi-Fi and Illumina sequencing platforms by Greif et al (2026. PMID: 41501640). The genome was assembled into 32 contigs, which represent complete chromosomes. The provided Fasta file also contains an additional contig corresponding to the maxicircle (mitochondrial genome) sequence. The Fasta files included in this dataset were downloaded from GenBank (assembly GCA_044048535.1; May 22, 2026). Additional information about the Dm28cT2T genome assembly and gene annotations may be accessed through the link: https://cruzi.pasteur.uy/
Burman, Nathaniel · Buyukyoruk, Murat · Wiegand, Tanner · et al.
4 files · 8.0 MB · fasta
This folder contains a multiple sequence alignment of Cas7 homologs in .fasta format, the domain-level annotations from PFAM and CasFinder, and an associated phylogenetic tree in .newick format.
Tek, Mümin İbrahim · Boyle, Brian · Normandeau, Eric · et al.
8.0 MB
Chromosome-scale genome assembly of four duckweed species: Spirodela polyrhiza, Lemna minuta, Lemna japonica and Lemna aequinoctialis x (hybrid between L. aequinoctialis and unresolved parental lineage)
This repository provides reformatted protein FASTA files derived from CNGB and MGnify/MetaPilot metagenomic catalogs obtained from human- and other vertebrate-associated microbiomes. To support downstream processing by bioinformatics pipelines, and in particular metaproteomic peptide identification from tandem mass spectrometry (MS/MS) data, all sequence headers were systematically reformatted into a standardized, UniProt-like syntax that embeds the original taxonomic and functional annotations directly within the header. 1. CNGB Catalogs For the CNGB microbiome protein catalogs ( https://db.cngb.org/microbiome ), the human gut Integrated Gene Catalogs (11M and 9M) were merged into a single database. To remove sequence redundancy, exact duplicate protein sequences were collapsed at 100% amino acid sequence identity using SeqKit . KEGG functional annotations available alongside the catalog were incorporated by converting the multi-accession TWINS.KEGG.catalog matrix into a mapping of gene IDs to KEGG Orthology (KO) terms. The KO terms were then converted into descriptive protein names using a reference KO annotation repository (v. 2026-02-01). The accompanying taxonomic annotations were not incorporated as, in most cases, their taxonomic nomenclature is now outdated. The same annotation, deduplication, and formatting workflows were applied to all CNGB microbiome catalogs for which functional annotations were available, including the mouse, pig, and rat gut microbiomes. The organism source descriptor ( OS= ) was customized in the UniProt-like headers to reflect the corresponding metagenomic source. 2. MGnify Catalogs The MGnify/MetaPilot microbiome protein catalogs consist of MGnify-derived collections of representative species-level metagenome-assembled genomes (MAGs), originally distributed through the Resources directory of the MetaPilot software repository ( https://github.com/cksakura/MetaPilot-Releases ). These catalogs were processed to integrate the taxonomic classification of each source MAG directly into the FASTA sequence headers while preserving the original functional protein descriptions. This release includes the following nine host-associated catalogs: Human-associated: Gut (4,744 species), Oral (452 species), Skin (579 species), Vaginal (280 species) Other vertebrate-associated: Mouse gut (2,847 species), Pig gut (1,376 species), Cow rumen (2,729 species), Sheep rumen (2,172 species), Chicken gut (1,322 species). Header Format All FASTA headers strictly follow the standard UniProt-like header format: >db|UniqueIdentifier|EntryName ProteinName OS=OrganismName PE=4 This format allows software tools that natively support UniProt databases to correctly parse and extract both taxonomic and functional metadata during database searching.
This dataset contains the genome sequence for Leishmania martiniquensis (LV760 strain). This genome sequence was de novo assembled using Oxford Nanopore and Illumina HiSeq sequencing methodologies by Almutairi et al (2021. PMID: 34137631). The genome was assembled into 42 contigs, of which 36 were categorized as chromosomes. In addition to the genome Fasta file, the Fasta files for 7967 protein-coding sequences and their corresponding protein sequences are also included in this dataset. The three Fasta files included in this dataset were downloaded from GenBank (assembly GCA_017916325.1; June 17, 2026).
This dataset contains the genome sequence for Leishmania tarentolae (Parrot Tar II). This genome sequence was de novo assembled using the PacBio RS II system by Goto et al (2020. PMID: 32439660). The genome was assembled into 179 chromosomes, of which 57 were assigned to specific chromosomes (1 to 36). In addition to the genome Fasta file, the Fasta files for 8703 protein-coding sequences and their corresponding protein sequences are also included in this dataset. The three Fasta files included in this dataset were downloaded from GenBank (assembly GCA_009731335.1; June 9, 2026).
DNA sequencing data: No mutations (SNPs or genetic rearrangements) were identified using the following tools Breseq Variant Report - v0.35. for the cells cultured with EVs, with and without Pmb. The reference sequence used to map the FASTQ files: Escherichia coli strain K-12 MG1655 ( E. coli reference genome), RefSeq ID: GCF_000005845.1 - genome assembly: ASM584v2.
Forero Junco, Laura Milena · Alanin, Katrine Wacenius Skov · M. Djurhuus, Amaru · et al.
8.0 MB
Dataset description This record contains seven FASTA files associated with the article "Bacteriophages Roam the Wheat Phyllosphere." The assemblies and derived viral sequence sets were generated from wheat phyllosphere viral metagenomes used to characterize bacteriophage diversity associated with winter wheat leaves. The dataset includes assemblies from flag leaf (FL) and penultimate leaf (PL) samples, both from native and multiple displacement amplification ( MDA ) virome preparations, as well as predicted viral contigs, representative virus operational taxonomic units ( vOTUs ), and two manually recovered complete viral genomes. The dataset includes: FL_amp_spades.fasta : amplified flag leaf virome assembly generated with SPAdes/metaSPAdes; 3,188 contigs . FL_spades.fasta : flag leaf virome assembly generated with SPAdes/metaSPAdes; 6,564 contigs . PL_amp_spades.fasta : amplified penultimate leaf virome assembly generated with SPAdes/metaSPAdes; 4,393 contigs . PL_spades.fasta : penultimate leaf virome assembly generated with SPAdes/metaSPAdes; 11,830 contigs . rmanually_recovered_complete_viral_genomes.fasta : two complete viral genomes manually recovered from subsampled assemblies: NODE_40, recovered from the 80% subsampling of PL, and APSE_wheat, recovered from the 10% subsampling of PL_mda; 2 complete viral genomes . viral_contigs.fasta : predicted viral contigs recovered from the wheat phyllosphere virome assemblies; 15,739 contigs . vOTUs.fasta : representative virus operational taxonomic units after dereplication/clustering; 876 vOTUs . Together, these files allow comparison of viral contigs recovered from wheat phyllosphere viromes across leaf types and amplification conditions. They support analyses of bacteriophage diversity, viral operational taxonomic units, phyllosphere-associated viral novelty, manually recovered complete viral genomes, and bacteriophage community structure in the wheat phyllosphere. Recommended citation Forero-Junco, L. M., Alanin, K. W. S., Djurhuus, A. M., Kot, W., Gobbi, A., & Hansen, L. H. (2022). Bacteriophages Roam the Wheat Phyllosphere. Viruses, 14 (2), 244. https://doi.org/10.3390/v14020244