Synthetic FASTA files used by the HSeeker benchmark suite (benchmarks/benchmark.py). Generated deterministically with fixed NumPy/Python random seeds documented in SEED_MANIFEST.json. Four sequence profiles: uniform (25 % each ACGT), ga_biased (90 % purine), ct_biased (90 % pyrimidine), realistic (~41 % GC). Two size tiers: small (~30 MB) and medium (~300 MB). Download via: python benchmarks/benchmark.py --from-zenodo
Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.
2,095 rows · 907 KB · fasta, tsv
A virus genome database representing 21,015 near-complete virus genomes collected from untargeted ultra-deep RNA/DNA combined sequencing of wastewater. Sequence data was provided by the CASPER consortium and raw data may be found on NCBI SRA under bioprojects PRJNA1247874 and PRJNA1198001. Data underwent read trimming, rRNA and human read removal, de novo assembly, and selection of high-quality viral contigs. Contigs were clustered at 95% identity and 85% query coverage to dereplicate. Chimera-checking required at least two independent assemblies of the same viral genome or presence of the genome in another reference database. Annotation made use of RdRpCATCH, geNomad, checkV, BLASTN against NCBI core-nt, and RNAVirHost. The RdRp fasta files contain representative RdRp sequences identified through homology to major RdRp reference databases and clustered at 90% sequence identity over 75% sequence coverage. Included sequences contain all three conserved RdRp motifs (A, B, and C) arranged in either the canonical ABC configuration or the permuted CAB configuration.
Hubbard, Alfred · Solares, Edwin · Hemming-Schroeder, Elizabeth
88 rows · 18 KB · fasta
These are the files needed to run the Broad Institute's malaria amplicon pipeline for the PvGAP Plasmodium vivax panel, described in detail here . They consist of FASTA files containing the forward and reverse primers and another FASTA file containing reference sequences for each target, derived from the PvP01 reference genome.
This dataset supports the findings of the manuscript "Root anatomical traits modulate the assembly and nitrogen transformation potential of root-associated microbiomes in a temperate steppe" (NPH-MS-2026-55667). It contains root traits data, bacterial 16S rRNA gene absolute abundances, functional genes relative abundances, DNA extraction metadata, and phylogenetic marker sequences for 37 plant species from a temperate steppe ecosystem. The dataset includes the following files: 1. root traits.csv - Root traits including average diameter (AD), specific root length (SRL), specific root area (SRA), root tissue density (RTD), root nitrogen content (RNC), root carbon content (RCC), carbon‑nitrogen ratio (RCN), cortex layer number (CLN), cortex thickness (CT), and the ratio of cortex thickness to root diameter (CTRD). The first column lists plant species names. 2. Absolute abundance of 16S rRNA gene.csv - Quantitative PCR (qPCR) derived absolute abundances of bacterial 16S rRNA gene copies (copies/ng DNA) across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 3. DNA extraction sample weight.csv - Fresh weight (grams) of root material used for DNA extraction for each sample, linked by SampleID to the abundance data. 4. DNA extraction concentration.csv - Qubit‑measured DNA concentrations (ng/μL) and the sample volume (μL) used for quality control, together with sample metadata. 5. 37species.fasta - DNA sequences of two chloroplast markers (matK and rbcL) for the 37 plant species included in the study. The sequences are in FASTA format with headers formatted as ">Species". These were used for host phylogeny construction and Pagel's λ analyses. 6. Quantitative PCR results of functional gene.csv - Quantitative PCR (qPCR) derived relative abundances of bacterial 16S rRNA gene and functional genes across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 7. README.md - A detailed description of each file, column headers, abbreviations, units, and any missing value codings (NA). All data are provided to ensure transparency and reproducibility of the analyses. For methodological details, please refer to the Materials and Methods section of the associated publication. These data are under embargo until the associated research article is published. After that date, they will be freely available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license. During the embargo period, the metadata (title, authors, abstract) and the DOI remain publicly visible, but the data files are not accessible. For access requests before the embargo expires, please contact the corresponding author.
The drug-target-interaction dataset is a combinatorial complex representing drug-target interactions and similarity relationships. The dataset is based on the work by Perlman et al., which combines multiple drug and gene similarity measures to predict drug-target interactions. Find the dataset details in AHORN .
This dataset contains 2134 unique malicious Python packages collected from PyPI, used for the evaluation of H2GLM (Hierarchical Heterogeneous Graph Learning framework enhanced by LLMs for Malicious package detection). Malicious samples are derived from a large-scale benchmark for malicious Python package detection, supplemented by packages collected from public security advisories and threat intelligence feeds. Deduplication was enforced via MD5 and fuzzy hashing. The archive malicious_2134.tar.gz contains the original .tar.gz distribution files as collected from PyPI. Intended for academic security research only.
# Data archive - Internal-state criticality in Bayesian-inverse-Bayesian inference **Paper.** *Internal-state criticality in Bayesian-inverse-Bayesian inference*, K. Sasai and Y.-P. Gunji (Physical Review Research, submitted). **Source repository.** <https://github.com/kazsasai/bayesian-inverse-bayesian-rps> **DOI.** `10.5281/zenodo.20533918`. --- ## Contents This deposit contains the raw simulation outputs underlying every data figure of the paper. Three archives are provided: two cover the main simulation data (*full reproducibility* vs *quick figure rebuild*), and one small archive holds the reinforcement-learning baseline-control data: | Archive | Size (compressed) | Contains | Use case | |---|---|---|---| | `paperA_data_full.tar.gz` | ~2.7 GB | Full simulation output tree (~16 GB uncompressed; 2713 files): all per-run JSONs and NPZs from `simulation/{reward_huge,nhand,reward_huge_v2,analyze_sharpness_plateau,reward}/data/` and `simulation_tie_mode_ablation/data/` | Independent re-analysis from raw outputs | | `paperA_data_figure_only.tar.gz` | ~1.1 GB | The 163 specific JSON/NPZ files actually read by `build_all.py` (~1.7 GB uncompressed) | Rebuild figures only | | `paperA_data_baseline_control.tar.gz` | ~21 MB | Pooled run-length arrays (`pnas_rl_comparison/data/baseline_dwells.npz`) for the RL-baseline control - WSLS, tabular Q-learning, and regret matching vs BIB; 40 seeds, T=2e5 - backing Fig. 4 (`fig_control_ab`) | Rebuild the RL-baseline control figure | The two main archives preserve the relative-path layout so that extracting either at `<repo>/data/` lets `build_all.py` find the data without further configuration. The baseline-control archive instead carries the `pnas_rl_comparison/data/...` path and extracts at the **repository root**. See **Reproducing the figures** below. Supporting files: * `MANIFEST_canonical.txt` - the in-repo data manifest (`BIB_Levy_v2/latex/figures/scripts/zenodo_data_manifest.txt`), listing each data tree, the figure(s) it feeds, and the generating script. * `figure_only_file_list.txt` - exhaustive 163-line list of relative paths inside `paperA_data_figure_only.tar.gz`, captured by auditing every `open()` call from a clean `build_all.py` run (and re-running with caches cleared so that no precomputed intermediate hid raw-data references). * `checksums.sha256` - SHA-256 of all three archives. ## Reproducing the figures Both tarballs preserve the same layout, so the workflow is identical: ```bash # 1. Clone the source repo git clone https://github.com/kazsasai/bayesian-inverse-bayesian-rps.git cd bayesian-inverse-bayesian-rps # 2. Get the data: pick ONE archive # (full = raw-output independent re-analysis; # figure-only = just enough to rebuild figures) mkdir -p data tar xzf /path/to/paperA_data_figure_only.tar.gz -C data # OR _full # 3. Install dependencies pip install numpy matplotlib powerlaw # 4. Rebuild figures python BIB_Levy_v2/latex/figures/scripts/build_all.py # (or run individual scripts: build_Fig3_universality.py, etc.) ``` Alternatively, point `PAPERA_DATA` at an extraction directory anywhere on disk: ```bash tar xzf paperA_data_figure_only.tar.gz -C /scratch/papera_data export PAPERA_DATA=/scratch/papera_data python BIB_Levy_v2/latex/figures/scripts/build_all.py ``` `figdata.py` in the source repo searches `$PAPERA_DATA`, then `<repo>/data/`, then the in-repo `simulation/` tree, in that order. ### RL-baseline control figure (Fig. 4) `paperA_data_baseline_control.tar.gz` carries the `pnas_rl_comparison/data/...` path, so extract it at the **repository root** (not `<repo>/data/`): ```bash tar xzf /path/to/paperA_data_baseline_control.tar.gz -C bayesian-inverse-bayesian-rps python pnas_si/figures/build_fig_control.py # -> fig_control_ab.{pdf,png} ``` The figure's BIB curves are read from the main data (the `reward_huge_*` `durations_bib-*` JSONs in the full / figure-only archive, via `$PAPERA_DATA`); the baseline curves come from the archive above. To regenerate the baseline data from scratch instead (deterministic, ~minutes): ```bash python pnas_rl_comparison/run_baseline_control.py # -> baseline_dwells.npz python pnas_rl_comparison/analyze_baseline_control.py ``` ## What `paperA_data_figure_only.tar.gz` excludes * The 17 G of per-run / per-step JSONs in the data trees that no current figure reads. * Intermediate caches (`fig*_ccdf_cache.json`) - these are regenerated by `build_Fig4_robustness.py` and `build_FigS2_nh_ccdf.py` on first run. * The small bundled inputs already shipped with the GitHub repo at `BIB_Levy_v2/latex/figures/scripts/data/` (`scheme_summary.csv`, `bo_tournament_results.json`, `sigma_*_rs_bib-bib.json`, `data_ivb_{equil,biased}.npz`). The build scripts read these straight from the repo. ## Verifying integrity ```bash shasum -a 256 -c checksums.sha256 ``` ## Citation If you use these data, please cite both the paper (forthcoming) and this Zenodo record. The repository's `README` is updated with the final citation on publication. ## License Data are released under CC-BY-4.0 (deposit metadata sets this on Zenodo). Source code in the GitHub repository is under its own LICENSE file.
Using gene-level and species-level shotgun metagenomics, we provide the first characterization of the rural, Bolivian microbiome; we identified microbial genes which strongly correlate (rho>0.45) with arsenic in urine, and that overall contribute to substantiate that the gut microbiome helps tolerate arsenic via a evict-out-of-house mechanism. Mediation analysis, followed by phylogenetic investigation of metagenomic-assembled genomes, further strengthens this observation. This study elucidates the role of the microbiome in helping to tolerate arsenic-rich environments, and paves the way for probiotic interventions that may mitigate the effects of this toxic metal. This repository contains 2,478 HQ MAGs from the MicroToxBol cohort in fasta format and a descriptive table comprising taxonomic annotation, quality-checks, and coverage estimation.
General description Compiled dataset of transcriptome assemblies, transcriptome annotation and expression quantification for 27 Amaryllidoideae species and 4 hybrid cultivars: Amaryllis belladonna , Clivia miniata , Crinum asiaticum , Crinum x powellii , Galanthus elwesii , Galanthus sp., Hippeastrum cv. Blossom Peacock, Hippeastrum cv. Jewel, Hippeastrum cv. Royal Velvet, Hippeastrum striatum , Hippeastrum vittatum , Leucojum aestivum , Lycoris aurea , Lycoris chinensis , Lycoris incarnata , Lycoris longituba , Lycoris radiata , Lycoris sprengeri , Narcissus aff pseudonarcissus , Narcissus papyraceus , Narcissus pseudonarcissus , Narcissus cv. Tête-à-Tête, Narcissus tazetta , Narcissus viridiflorus , Phycella aff cyrtanthoides , Rhodophiala pratensis , Scadoxus multiflorus , Traubia modesta , Zephyranthes candida , Zephyranthes carinata , Zephyranthes treatiae. Data includes transcriptome assemblies, Transdecoder predictions of peptide sequences (along with predicted coding seqeunces, GFF3 and BED files), expression quantification (count and TPM matrices for genes and trinity isoforms obtained with Kallisto and from Salmon), and annotation results (from EggNOG, Pfam, Uniprot Swissprot and Rfam), as well as signal peptide and transmemberane domain predictions (from SignalP and TmHMM). Annotations were compiled into a report for each species using Trinotate. For Lycoris aurea and Narcissus cv. Tête-à-Tête, there are two assemblies and corresponding files: Lycoris_aurea_PB and Lycoris_aurea_TH, and Narcissus_TaT_PB Narcissus_TaT_TH. PB assemblies were constructed solely with long-read sequencing data (PacBio, PB); while TH assemblies were constructed with short reads using Trinity, with long-read assembly being used for scaffolding step of Trinity (Trinity Hybrid, TH). Expression quantification was published on NCBI GEO (accessions GSE329951 , GSE329957 , GSE330014 and GSE331457 ). File descriptions All unitigs and predicted protein sequences are prefixed with an acronym to identify the species: Species Acronym NCBI TSA accession Amaryllis belladonna Ambel deposited on GSE331457 Clivia miniata Clmin DBNKRK000000000 Crinum asiaticum Crasi DBNIJL000000000 Crinum x powellii Crpow DBNKRS000000000 Galanthus elwesii Gaelw DBNKRO000000000 Galanthus sp. Gasp DBNKRR000000000 Hippeastrum cv. Blossom Peacock HispBP DBNMYE000000000 Hippeastrum cv. Jewel HispJW DBNMYC000000000 Hippeastrum cv. Royal Velvet HispRV DBNMYD000000000 Hippeastrum striatum Histr DBNIJG000000000 Hippeastrum vittatum Hivit DBNKRF000000000 Leucojum aestivum Leaes DBNIJJ000000000 Lycoris aurea (PB) Lyaur DBNFTY000000000 Lycoris aurea (TH) Lyaur DBNKRI000000000 Lycoris chinensis Lychi DBNKRE000000000 Lycoris incarnata Lyinc DBNKRD000000000 Lycoris longituba Lylon DBNKRG000000000 Lycoris radiata Lyrad deposited on GSE331457 Lycoris sprengeri Lyspr DBNKRL000000000 Narcissus aff pseudonarcissus Naafps DBNKRP000000000 Narcissus papyraceus Napap DBNKRN000000000 Narcissus pseudonarcissus Napse DBNIJK000000000 Narcissus Tête-à-Tête (PB) NaspPB DBNNFO000000000 Narcissus Tête-à-Tête (TH) NaspTH DBNUFN000000000 Narcissus tazetta Nataz DBNPME000000000 Narcissus viridiflorus Navir deposited on GSE331457 Phycella sp. Phsp deposited on GSE331457 Rhodophiala pratensis Rhpra deposited on GSE331457 Scadoxus multiflorus Scmul DBNIJH000000000 Traubia modesta Trmod deposited on GSE331457 Zephyranthes candida Zecan DBNKRH000000000 Zephyranthes carinata Zecar DBNIJI000000000 Zephyranthes treatiae Zetre deposited on GSE331457 Files Files are organized by type of analysis/data, meaning all expression quantification data obtained with Kallisto are compressed into a single file, all results from BlastP are in the same file, etc. Compressed file name Individual file type Content Blastp_Uniprot.tar.gz Tabular 33 files (1 per assembly) generated with BLASTP against Uniprot SwissProt release 2024_04. Files in blast output format 6 (standard columns). EggNOG_Emapper.tar.gz Tabular 33 tabular files (1 per assembly). Generated with eggnog.emapper (annotations file format described in eggnog.emapper's wiki ) Expression_Kallisto.tar.gz Tabular 128 files (4 per assembly, 2 for long-read assemblies): read counts and TPM values for both trinity "genes" and trinity "isoforms"). Row names are "gene"/ "isoform" IDs, column names are SRA run IDs. Expression_Salmon.tar.gz Tabular 128 files (4 per assembly, 2 for long-read assemblies): read counts and TPM values for both trinity "genes" and trinity "isoforms"). Row names are "gene"/ "isoform" IDs, column names are SRA run IDs. Final_assemblies.tar.gz Fasta 33 files (1 per assembly). Assemblies generated in this study, after contamination and expression filtering. HMMScan_PfamA.tar.gz Domain hits table 33 files (1 per assembly). Generated with hmmscan (option ‑‑domtblout, domain hits table, explained in hmmer's user guide [119] ) against the Pfam-A database. Infernal_Rfam.tar.gz Target hits table format 2 33 files (1 per assembly) generated with Infernal's cmscan (using the Trinotate wrapper) against the Rfam database. Table format 2 described in Infernal's user guide section 6 ). Metadata_studies.tar.gz Tabular (semi-colon separated columns) 31 tabular files (one per species/cultivar) indicating: Bioproject ID, SRA run ID, Sample name, Biosample ID, tissue, genotype (cultivar, when specified), treatment, batch, original publication citation, and DOI of original publication. Signalp6.tar.gz Tabular or GFF3 99 files (3 per assembly: prediction_results.txt, output.gff3 and region_output.gff3) generated with SignalP6. TmHMM2.tar.gz Tabular 33 files (1 per assembly) generated with TmHMM2 (short format, described in the guide tab of https://services.healthtech.dtu.dk/services/TMHMM-2.0/ Transcriptomes_unfiltered.tar.gz Fasta 33 files (1 per assembly). Unfiltered (prior to contamination and expression screening) assemblies generated in this study. Transdecoder_bed.tar.gz BED 33 BED files (1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam) Transdecoder_cds.tar.gz Fasta 33 fasta files (1 per assembly) of predicted coding sequences generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam) Transdecoder_gff3.tar.gz GFF3 33 GFF3 files (1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam). Transdecoder_proteomes.tar.gz Fasta 33 predicted proteome files (.pep, 1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam).
R. Panicker, Vyshakh · J. Smug, Bogna · Klein-Sousa, Victor · et al.
1 files · 8.0 MB · gzip
Data deposit for: Panicker VR, Smug BJ, Klein-Sousa V, Enright MC, Taylor NMI, Drulis-Kawa Z, Mostowy RJ. "Structural modularity of receptor-binding proteins underlies host-range strategy diversification in Klebsiella pneumoniae phages." bioRxiv 2026. https://doi.org/10.64898/2026.05.12.724579 This archive contains all large binary data files required to reproduce the analyses and figures in the paper above. The associated code is available at: https://github.com/VyshakhRP/RBP-div-hostrange Contents: - 01_raw/02_assemblies/ - Raw lysate assemblies for unpublished Klebsiella pneumoniae phages - 01_raw/03_host_genomes/ - Klebsiella pneumoniae host genome sequences (.fasta) - 01_raw/04_phage_genomes/ - Phage genome sequences (.fasta, .gb) - 02_intermediate/09_af3_predictions/ - AlphaFold 3 structure predictions for all receptor-binding proteins (.cif) - 02_intermediate/14_rbps-ecods/ecod.develop288.domains.txt - ECOD domain database (v20230309, develop288) - 02_intermediate/14_rbps-ecods/ecod.develop288.fasta.txt - ECOD domain FASTA (v20230309, develop288) - 05_output/01_genomes/ - Processed phage genome files - 05_output/03_proteins/ - Protein FASTAs and receptor-binding protein structures Usage: Download the archive, extract it, and place each subdirectory into the corresponding location in the cloned GitHub repository. Full instructions are provided in the repository README under Data Availability. Note: The ECOD database files can alternatively be downloaded directly from http://prodata.swmed.edu/ecod/distributions/ (v20230309 / develop288).
Requena Rolanía, Jose María · Greif, Gonzalo · ROBELLO, CARLOS
1 files · 8.0 MB · fasta
This dataset contains the genome sequence for Trypanosoma cruzi (strain Dm28c). This genome sequence was de novo assembled using PacBio Hi-Fi and Illumina sequencing platforms by Greif et al (2026. PMID: 41501640). The genome was assembled into 32 contigs, which represent complete chromosomes. The provided Fasta file also contains an additional contig corresponding to the maxicircle (mitochondrial genome) sequence. The Fasta files included in this dataset were downloaded from GenBank (assembly GCA_044048535.1; May 22, 2026). Additional information about the Dm28cT2T genome assembly and gene annotations may be accessed through the link: https://cruzi.pasteur.uy/
These files are downsampled WGS fastq files (250k paired-end reads each) of a fungal pathogen ( Ascochyta rabiei ), generated on MGI DNSeq-T7. The files are intended to be used directly as test datasets in PopFun - a Nextflow pipeline for variant calling in fungal genomes.
Zhang, Zirui · Jones, Ashley · Schwessinger, Benjamin · et al.
1 files · 8.0 MB · gzip
This dataset contains annotation supporting files for the chromosome-scale, haplotype-resolved genome assembly of the sexually deceptive Australian orchid Chiloglottis trapeziformis. The archive includes haplotype-specific BRAKER gene annotation files, predicted coding sequences, predicted protein sequences, functional annotation results, FASTA index files, a file manifest, and SHA256 checksums for haplotype 1 and haplotype 2. This record is intended to support the manuscript "Inter-haplotype inversions and repeat expansion in the sexually deceptive orchid Chiloglottis trapeziformis".
Burman, Nathaniel · Buyukyoruk, Murat · Wiegand, Tanner · et al.
4 files · 8.0 MB · fasta
This folder contains a multiple sequence alignment of Cas7 homologs in .fasta format, the domain-level annotations from PFAM and CasFinder, and an associated phylogenetic tree in .newick format.
This data record contains all input data and scripts to rerun the simulations and it includes the output structures. The dataset contains the raw and processed data used to create figures 5I, 5J of the assocated manuscript.
This dataset contains the imputed SNP genotypes of 406 bread wheat accessions in VCF format, used for haplotype-based GWAS analysis of root architecture.
README - Data for "Genomic constraint and hypervariability in tetraploid potatoes" Trine Aalborg, May 2026 The data applied in the study includes phenotypic and genotypic information on the MASPOT panel (768 F1 progeny of an 18-parent diallel cross - property of Danespo A/S). The genotypic data was generated using genotyping-by-sequencing technology as described in the paper. Following genotype calling and filtration (5-60x read depth, < 50 % missing rate, > 1 % MAF), the total SNP set includes 151,164 biallelic SNPs. Coordinates of these SNPs relative to the DMv6.1 potato reference genome are provided. In addition to genotypes across the clones, the estimated GERP score of that SNP from (Wu et al., 2023) is reported. The manuscript analyses only consider markers with reliable GERP scores (MSA alignment depth > 50, and neutral score > 2), which corresponded to 97,815 of the total 151,164 biallelic SNPs. SNPeff annotations of the markers (based on the DMv6.1 reference genome) are also appended. The phenotypes were collected across 1-2 field trials, depending on the traits, and includes a minimum of two replicates per clone from a randomized block design. There are phenotypes for eight traits: dry matter content [%], yield (hkg/ha), senescence [1-9], flesh color [1-9], tubers/plant, tuber length [mm], tuber diameter [mm], and tuber size [mm^3]. Metadata includes phenotyping year and block location of the plot as well as pedigree of the diallel offspring. File descriptions: gt_MASPOT.csv - .csv file of the non-imputed genotypic data of 151,164 SNPs for the 768 F1 clones (those with GERP scores). Columns 1-3 are SNP coordinates and SNP IDs. Column names from column 4 and onwards are clone IDs. gt_MASPOT_imputed.csv - .csv file holding the imputed (random forest, missRanger algorithm) genotypic data of 151,164 SNPs for the 768 F1 clones. gt_MASPOT_recoded.csv - .csv file of the recoded, imputed genotypic data of 151,164 SNPs for the 768 F1 clones. The SNPs are recoded from original MASPOT ref/alt allele (based on AF in the MASPOT panel) to the alternative allele = the derived allele in the 100-Solanaceae panel. pt_MASPOT.csv - .csv file holding the phenotypic data (eight traits) of the 768 F1 clones (Clone_ID). The number of observations varies across traits. Also including metadata: year of phenotyping (Year), block number (Block, Line_in_block), clone parents (Mother, Father, Family). GERP_MASPOT.csv - .csv file holding the GERP scores (including alignment depth, neutral scores, and a marker annotation based on GERP score thresholds (deleterious, neutral, hypervariable, or low quality)), SNPeff annotations, and the derived allele in the 100-Solanaceae panel from (Wu et al., 2023) [MASPOT_Alt_Allele_Is_Sol_Derived_Allele - used for recoding of the genotypes] of the MASPOT SNPs with GERP scores. snps.MASPOT_F1.vcf.gz - zipped .vcf file of the GBS MASPOT genotypic data (both discrete genotype calls and allele frequencies) called to the DMv6.1 potato reference genome. Filtered to read depth 5x, MQ > 30. A total of 160,920 biallelic SNPs. Includes the 768 F1 progeny analyzed. Literature: Wu, Y., Li, D., Hu, Y., Li, H., Ramstein, G. P., Zhou, S., et al. (2023). Phylogenomic discovery of deleterious mutations facilitates hybrid potato breeding. Cell 186, 2313-2328.e15. doi: 10.1016/j.cell.2023.04.008
Expected results from running comprehensive tests using zol v1.6.19. Script for running tests also included (but is primarily updated/tracked on zol's GitHub repo). The latest testing dataset is also included, correcting prior issues with corrupted FASTA files of Aspergillus genomes.
Jiménez-Blanco, Albert · López-Villellas, Lorién · Moure, Juan Carlos · et al.
1 files · 8.0 MB · gzip
Datasets used in the "Theseus: Fast and Optimal Affine-Gap Sequence-to-Graph Alignment" paper. The data is compressed into the theseus_datasets.tar.gz file. This file includes the datasets for the two experiments on the paper: MSA datasets: File Size Source hiv_100.fasta 100 sequences, each of approximately 10Kbp https://github.com/niemasd/ViralMSA/blob/master/example/example_hiv.fas mtb_benchmark_50kbp_shortened.fna 342 sequences, each of approximately 50Kbp Derived from a set of 342 RefSeq-complete whole-genome assemblies of Mycobacterium tuberculosis genomes. Sequences start at the dnaA gene and are 50Kbp long. mtb_benchmark_250kbp_trmB.fna 342 sequences, each of approximately 250Kbp Derived from a set of 342 RefSeq-complete whole-genome assemblies of Mycobacterium tuberculosis genomes. Sequences start at the dnaA gene and are truncated at the trmB gene. mtb_benchmark_500kbp_thiE.fna 342 sequences, each of approximately 500Kbp Derived from a set of 342 RefSeq-complete whole-genome assemblies of Mycobacterium tuberculosis genomes. Sequences start at the dnaA gene and are truncated at the thiE gene. mtb_benchmark_1Mbp_gltA2.fna 342 sequences, each of approximately 1Mbp Derived from a set of 342 RefSeq-complete whole-genome assemblies of Mycobacterium tuberculosis genomes. Sequences start at the dnaA gene and are truncated at the gltA2 gene. covid_19_complete.fasta 2732 sequences, each of approximately ∼30 Kbp in length 2732 GenBank-complete SARS-CoV-2 genome assemblies. monkeypox_100_seq.fasta 100 sequences, each of approximately ∼200 Kbp in length 100 RefSeq-complete whole genome assemblies of Monkey pox's virus. Sequence-to-graph /pangenome read mapping datasets: File Size Source SRR062634_1.filt_REDUCED.fasta 250K sequences of length 100bp Human Pangenome Reference Consortium 211109_M024_V350038332_L01_HUMuarfR092940-606_1_REDUCED.fasta 250K sequences of length 150bp NIST Genome in a Bottle (GIAB) project D1_S1_L001_R1_001_REDUCED.fasta 250K sequences of length 250bp NIST Genome in a Bottle (GIAB) project This dataset is derived from original data produced by third parties, as detailed above. All rights to the original data remain with the original authors or copyright holders. Users are responsible for ensuring compliance with the licensing terms of the original data sources. Sequence-to-cyclic-graph datasets: These datasets have been synthetically generated to test the ability of Theseus to align against cyclic reference graphs, comparing it to the graph unfolding strategy used by vg map. The synthetic cyclic graphs with a controlled structure. For that, we create backbone graphs with N in {10, 100, 1000} nodes, with node sequences of average length 50 bp, and connect each node to its successor to ensure graph connectivity. Then, we add a bounded number of random edges per node. These random edges can connect to nearby forward nodes or to previously created nodes, introducing cycles into the graph. For each graph size, we generate queries of length 100, 150, and 250 bp. We have nine pairs of reference graph and queries. This happens because we have three graph sizes, with N in {10, 100, 1000} nodes, and 3 query sizes, of 100, 150, and 250 base pairs. Code snapshot: This dataset also contains a snapshot of the code use to conduct the experiments in the manuscript, corresponding to Theseus v0.1 on Github.