HSeeker Benchmark FASTA Datasets
HSeeker Contributors
7 files · 7.5 MB · gzip
hybrid · semantic + lexical · 149 datasets ranked · 0.73s
HSeeker Contributors
7 files · 7.5 MB · gzip
Synthetic FASTA files used by the HSeeker benchmark suite (benchmarks/benchmark.py). Generated deterministically with fixed NumPy/Python random seeds documented in SEED_MANIFEST.json. Four sequence profiles: uniform (25 % each ACGT), ga_biased (90 % purine), ct_biased (90 % pyrimidine), realistic (~41 % GC). Two size tiers: small (~30 MB) and medium (~300 MB). Download via: python benchmarks/benchmark.py --from-zenodo
Computational Network Science
1 files · 415 KB · gzip
The drug-target-interaction dataset is a combinatorial complex representing drug-target interactions and similarity relationships. The dataset is based on the work by Perlman et al., which combines multiple drug and gene similarity measures to predict drug-target interactions. Find the dataset details in AHORN .
Altenhoff, Adrian
6 files · 100 MB · hdf5
OMAmer - tree-driven and alignment-free protein assignment to subfamilies OMAmer is an alignment-free protein family assignment method designed to avoid overly specific subfamily predictions and to scale efficiently to phylogenomic databases containing thousands of genomes. It relies on an innovative approach that uses evolutionarily informed k-mers for alignment-free mapping to ancestral protein subfamilies. This dataset provides precomputed OMAmer databases derived from the Hierarchical Orthologous Groups in the OMA Browser . We aim to update these databases with every new OMA Browser release. Each OMAmer database is built using the latest version of the OMAmer package available at the time of the corresponding OMA Browser release. The dataset includes databases for different subsets of the species taxonomy. In most cases, we recommend using the LUCA.h5 database, which contains information from all species in the OMA database. The subset-specific databases are mainly useful when disk space is limited. The release May2026 is based on the OMA Browser release May 2026 which comprises 2983 species. We used OMAmer version 2.1.0 to build these databases.
Anonymous
1 files · 8.0 MB · gzip
This dataset contains 2134 unique malicious Python packages collected from PyPI, used for the evaluation of H2GLM (Hierarchical Heterogeneous Graph Learning framework enhanced by LLMs for Malicious package detection). Malicious samples are derived from a large-scale benchmark for malicious Python package detection, supplemented by packages collected from public security advisories and threat intelligence feeds. Deduplication was enforced via MD5 and fuzzy hashing. The archive malicious_2134.tar.gz contains the original .tar.gz distribution files as collected from PyPI. Intended for academic security research only.
Sasai, Kazuto · Yukio-Pegio, Gunji
7 files · 8.0 MB · gzip
# Data archive - Internal-state criticality in Bayesian-inverse-Bayesian inference **Paper.** *Internal-state criticality in Bayesian-inverse-Bayesian inference*, K. Sasai and Y.-P. Gunji (Physical Review Research, submitted). **Source repository.** <https://github.com/kazsasai/bayesian-inverse-bayesian-rps> **DOI.** `10.5281/zenodo.20533918`. --- ## Contents This deposit contains the raw simulation outputs underlying every data figure of the paper. Three archives are provided: two cover the main simulation data (*full reproducibility* vs *quick figure rebuild*), and one small archive holds the reinforcement-learning baseline-control data: | Archive | Size (compressed) | Contains | Use case | |---|---|---|---| | `paperA_data_full.tar.gz` | ~2.7 GB | Full simulation output tree (~16 GB uncompressed; 2713 files): all per-run JSONs and NPZs from `simulation/{reward_huge,nhand,reward_huge_v2,analyze_sharpness_plateau,reward}/data/` and `simulation_tie_mode_ablation/data/` | Independent re-analysis from raw outputs | | `paperA_data_figure_only.tar.gz` | ~1.1 GB | The 163 specific JSON/NPZ files actually read by `build_all.py` (~1.7 GB uncompressed) | Rebuild figures only | | `paperA_data_baseline_control.tar.gz` | ~21 MB | Pooled run-length arrays (`pnas_rl_comparison/data/baseline_dwells.npz`) for the RL-baseline control - WSLS, tabular Q-learning, and regret matching vs BIB; 40 seeds, T=2e5 - backing Fig. 4 (`fig_control_ab`) | Rebuild the RL-baseline control figure | The two main archives preserve the relative-path layout so that extracting either at `<repo>/data/` lets `build_all.py` find the data without further configuration. The baseline-control archive instead carries the `pnas_rl_comparison/data/...` path and extracts at the **repository root**. See **Reproducing the figures** below. Supporting files: * `MANIFEST_canonical.txt` - the in-repo data manifest (`BIB_Levy_v2/latex/figures/scripts/zenodo_data_manifest.txt`), listing each data tree, the figure(s) it feeds, and the generating script. * `figure_only_file_list.txt` - exhaustive 163-line list of relative paths inside `paperA_data_figure_only.tar.gz`, captured by auditing every `open()` call from a clean `build_all.py` run (and re-running with caches cleared so that no precomputed intermediate hid raw-data references). * `checksums.sha256` - SHA-256 of all three archives. ## Reproducing the figures Both tarballs preserve the same layout, so the workflow is identical: ```bash # 1. Clone the source repo git clone https://github.com/kazsasai/bayesian-inverse-bayesian-rps.git cd bayesian-inverse-bayesian-rps # 2. Get the data: pick ONE archive # (full = raw-output independent re-analysis; # figure-only = just enough to rebuild figures) mkdir -p data tar xzf /path/to/paperA_data_figure_only.tar.gz -C data # OR _full # 3. Install dependencies pip install numpy matplotlib powerlaw # 4. Rebuild figures python BIB_Levy_v2/latex/figures/scripts/build_all.py # (or run individual scripts: build_Fig3_universality.py, etc.) ``` Alternatively, point `PAPERA_DATA` at an extraction directory anywhere on disk: ```bash tar xzf paperA_data_figure_only.tar.gz -C /scratch/papera_data export PAPERA_DATA=/scratch/papera_data python BIB_Levy_v2/latex/figures/scripts/build_all.py ``` `figdata.py` in the source repo searches `$PAPERA_DATA`, then `<repo>/data/`, then the in-repo `simulation/` tree, in that order. ### RL-baseline control figure (Fig. 4) `paperA_data_baseline_control.tar.gz` carries the `pnas_rl_comparison/data/...` path, so extract it at the **repository root** (not `<repo>/data/`): ```bash tar xzf /path/to/paperA_data_baseline_control.tar.gz -C bayesian-inverse-bayesian-rps python pnas_si/figures/build_fig_control.py # -> fig_control_ab.{pdf,png} ``` The figure's BIB curves are read from the main data (the `reward_huge_*` `durations_bib-*` JSONs in the full / figure-only archive, via `$PAPERA_DATA`); the baseline curves come from the archive above. To regenerate the baseline data from scratch instead (deterministic, ~minutes): ```bash python pnas_rl_comparison/run_baseline_control.py # -> baseline_dwells.npz python pnas_rl_comparison/analyze_baseline_control.py ``` ## What `paperA_data_figure_only.tar.gz` excludes * The 17 G of per-run / per-step JSONs in the data trees that no current figure reads. * Intermediate caches (`fig*_ccdf_cache.json`) - these are regenerated by `build_Fig4_robustness.py` and `build_FigS2_nh_ccdf.py` on first run. * The small bundled inputs already shipped with the GitHub repo at `BIB_Levy_v2/latex/figures/scripts/data/` (`scheme_summary.csv`, `bo_tournament_results.json`, `sigma_*_rs_bib-bib.json`, `data_ivb_{equil,biased}.npz`). The build scripts read these straight from the repo. ## Verifying integrity ```bash shasum -a 256 -c checksums.sha256 ``` ## Citation If you use these data, please cite both the paper (forthcoming) and this Zenodo record. The repository's `README` is updated with the final citation on publication. ## License Data are released under CC-BY-4.0 (deposit metadata sets this on Zenodo). Source code in the GitHub repository is under its own LICENSE file.
Manghi, Paolo
2 files · 8.0 MB · gzip, tsv
Using gene-level and species-level shotgun metagenomics, we provide the first characterization of the rural, Bolivian microbiome; we identified microbial genes which strongly correlate (rho>0.45) with arsenic in urine, and that overall contribute to substantiate that the gut microbiome helps tolerate arsenic via a evict-out-of-house mechanism. Mediation analysis, followed by phylogenetic investigation of metagenomic-assembled genomes, further strengthens this observation. This study elucidates the role of the microbiome in helping to tolerate arsenic-rich environments, and paves the way for probiotic interventions that may mitigate the effects of this toxic metal. This repository contains 2,478 HQ MAGs from the MicroToxBol cohort in fasta format and a descriptive table comprising taxonomic annotation, quality-checks, and coverage estimation.
Goncalves dos Santos, Karen Cristine · Desgagné-Penix, Isabel · Merindol, Natacha
16 files · 8.0 MB · gzip
General description Compiled dataset of transcriptome assemblies, transcriptome annotation and expression quantification for 27 Amaryllidoideae species and 4 hybrid cultivars: Amaryllis belladonna , Clivia miniata , Crinum asiaticum , Crinum x powellii , Galanthus elwesii , Galanthus sp., Hippeastrum cv. Blossom Peacock, Hippeastrum cv. Jewel, Hippeastrum cv. Royal Velvet, Hippeastrum striatum , Hippeastrum vittatum , Leucojum aestivum , Lycoris aurea , Lycoris chinensis , Lycoris incarnata , Lycoris longituba , Lycoris radiata , Lycoris sprengeri , Narcissus aff pseudonarcissus , Narcissus papyraceus , Narcissus pseudonarcissus , Narcissus cv. Tête-à-Tête, Narcissus tazetta , Narcissus viridiflorus , Phycella aff cyrtanthoides , Rhodophiala pratensis , Scadoxus multiflorus , Traubia modesta , Zephyranthes candida , Zephyranthes carinata , Zephyranthes treatiae. Data includes transcriptome assemblies, Transdecoder predictions of peptide sequences (along with predicted coding seqeunces, GFF3 and BED files), expression quantification (count and TPM matrices for genes and trinity isoforms obtained with Kallisto and from Salmon), and annotation results (from EggNOG, Pfam, Uniprot Swissprot and Rfam), as well as signal peptide and transmemberane domain predictions (from SignalP and TmHMM). Annotations were compiled into a report for each species using Trinotate. For Lycoris aurea and Narcissus cv. Tête-à-Tête, there are two assemblies and corresponding files: Lycoris_aurea_PB and Lycoris_aurea_TH, and Narcissus_TaT_PB Narcissus_TaT_TH. PB assemblies were constructed solely with long-read sequencing data (PacBio, PB); while TH assemblies were constructed with short reads using Trinity, with long-read assembly being used for scaffolding step of Trinity (Trinity Hybrid, TH). Expression quantification was published on NCBI GEO (accessions GSE329951 , GSE329957 , GSE330014 and GSE331457 ). File descriptions All unitigs and predicted protein sequences are prefixed with an acronym to identify the species: Species Acronym NCBI TSA accession Amaryllis belladonna Ambel deposited on GSE331457 Clivia miniata Clmin DBNKRK000000000 Crinum asiaticum Crasi DBNIJL000000000 Crinum x powellii Crpow DBNKRS000000000 Galanthus elwesii Gaelw DBNKRO000000000 Galanthus sp. Gasp DBNKRR000000000 Hippeastrum cv. Blossom Peacock HispBP DBNMYE000000000 Hippeastrum cv. Jewel HispJW DBNMYC000000000 Hippeastrum cv. Royal Velvet HispRV DBNMYD000000000 Hippeastrum striatum Histr DBNIJG000000000 Hippeastrum vittatum Hivit DBNKRF000000000 Leucojum aestivum Leaes DBNIJJ000000000 Lycoris aurea (PB) Lyaur DBNFTY000000000 Lycoris aurea (TH) Lyaur DBNKRI000000000 Lycoris chinensis Lychi DBNKRE000000000 Lycoris incarnata Lyinc DBNKRD000000000 Lycoris longituba Lylon DBNKRG000000000 Lycoris radiata Lyrad deposited on GSE331457 Lycoris sprengeri Lyspr DBNKRL000000000 Narcissus aff pseudonarcissus Naafps DBNKRP000000000 Narcissus papyraceus Napap DBNKRN000000000 Narcissus pseudonarcissus Napse DBNIJK000000000 Narcissus Tête-à-Tête (PB) NaspPB DBNNFO000000000 Narcissus Tête-à-Tête (TH) NaspTH DBNUFN000000000 Narcissus tazetta Nataz DBNPME000000000 Narcissus viridiflorus Navir deposited on GSE331457 Phycella sp. Phsp deposited on GSE331457 Rhodophiala pratensis Rhpra deposited on GSE331457 Scadoxus multiflorus Scmul DBNIJH000000000 Traubia modesta Trmod deposited on GSE331457 Zephyranthes candida Zecan DBNKRH000000000 Zephyranthes carinata Zecar DBNIJI000000000 Zephyranthes treatiae Zetre deposited on GSE331457 Files Files are organized by type of analysis/data, meaning all expression quantification data obtained with Kallisto are compressed into a single file, all results from BlastP are in the same file, etc. Compressed file name Individual file type Content Blastp_Uniprot.tar.gz Tabular 33 files (1 per assembly) generated with BLASTP against Uniprot SwissProt release 2024_04. Files in blast output format 6 (standard columns). EggNOG_Emapper.tar.gz Tabular 33 tabular files (1 per assembly). Generated with eggnog.emapper (annotations file format described in eggnog.emapper's wiki ) Expression_Kallisto.tar.gz Tabular 128 files (4 per assembly, 2 for long-read assemblies): read counts and TPM values for both trinity "genes" and trinity "isoforms"). Row names are "gene"/ "isoform" IDs, column names are SRA run IDs. Expression_Salmon.tar.gz Tabular 128 files (4 per assembly, 2 for long-read assemblies): read counts and TPM values for both trinity "genes" and trinity "isoforms"). Row names are "gene"/ "isoform" IDs, column names are SRA run IDs. Final_assemblies.tar.gz Fasta 33 files (1 per assembly). Assemblies generated in this study, after contamination and expression filtering. HMMScan_PfamA.tar.gz Domain hits table 33 files (1 per assembly). Generated with hmmscan (option ‑‑domtblout, domain hits table, explained in hmmer's user guide [119] ) against the Pfam-A database. Infernal_Rfam.tar.gz Target hits table format 2 33 files (1 per assembly) generated with Infernal's cmscan (using the Trinotate wrapper) against the Rfam database. Table format 2 described in Infernal's user guide section 6 ). Metadata_studies.tar.gz Tabular (semi-colon separated columns) 31 tabular files (one per species/cultivar) indicating: Bioproject ID, SRA run ID, Sample name, Biosample ID, tissue, genotype (cultivar, when specified), treatment, batch, original publication citation, and DOI of original publication. Signalp6.tar.gz Tabular or GFF3 99 files (3 per assembly: prediction_results.txt, output.gff3 and region_output.gff3) generated with SignalP6. TmHMM2.tar.gz Tabular 33 files (1 per assembly) generated with TmHMM2 (short format, described in the guide tab of https://services.healthtech.dtu.dk/services/TMHMM-2.0/ Transcriptomes_unfiltered.tar.gz Fasta 33 files (1 per assembly). Unfiltered (prior to contamination and expression screening) assemblies generated in this study. Transdecoder_bed.tar.gz BED 33 BED files (1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam) Transdecoder_cds.tar.gz Fasta 33 fasta files (1 per assembly) of predicted coding sequences generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam) Transdecoder_gff3.tar.gz GFF3 33 GFF3 files (1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam). Transdecoder_proteomes.tar.gz Fasta 33 predicted proteome files (.pep, 1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam).
R. Panicker, Vyshakh · J. Smug, Bogna · Klein-Sousa, Victor · et al.
1 files · 8.0 MB · gzip
Data deposit for: Panicker VR, Smug BJ, Klein-Sousa V, Enright MC, Taylor NMI, Drulis-Kawa Z, Mostowy RJ. "Structural modularity of receptor-binding proteins underlies host-range strategy diversification in Klebsiella pneumoniae phages." bioRxiv 2026. https://doi.org/10.64898/2026.05.12.724579 This archive contains all large binary data files required to reproduce the analyses and figures in the paper above. The associated code is available at: https://github.com/VyshakhRP/RBP-div-hostrange Contents: - 01_raw/02_assemblies/ - Raw lysate assemblies for unpublished Klebsiella pneumoniae phages - 01_raw/03_host_genomes/ - Klebsiella pneumoniae host genome sequences (.fasta) - 01_raw/04_phage_genomes/ - Phage genome sequences (.fasta, .gb) - 02_intermediate/09_af3_predictions/ - AlphaFold 3 structure predictions for all receptor-binding proteins (.cif) - 02_intermediate/14_rbps-ecods/ecod.develop288.domains.txt - ECOD domain database (v20230309, develop288) - 02_intermediate/14_rbps-ecods/ecod.develop288.fasta.txt - ECOD domain FASTA (v20230309, develop288) - 05_output/01_genomes/ - Processed phage genome files - 05_output/03_proteins/ - Protein FASTAs and receptor-binding protein structures Usage: Download the archive, extract it, and place each subdirectory into the corresponding location in the cloned GitHub repository. Full instructions are provided in the repository README under Data Availability. Note: The ECOD database files can alternatively be downloaded directly from http://prodata.swmed.edu/ecod/distributions/ (v20230309 / develop288).
Bar, Ido
9 files · 8.0 MB · gff, gzip
These files are downsampled WGS fastq files (250k paired-end reads each) of a fungal pathogen ( Ascochyta rabiei ), generated on MGI DNSeq-T7. The files are intended to be used directly as test datasets in PopFun - a Nextflow pipeline for variant calling in fungal genomes.
Zhang, Zirui · Jones, Ashley · Schwessinger, Benjamin · et al.
1 files · 8.0 MB · gzip
This dataset contains annotation supporting files for the chromosome-scale, haplotype-resolved genome assembly of the sexually deceptive Australian orchid Chiloglottis trapeziformis. The archive includes haplotype-specific BRAKER gene annotation files, predicted coding sequences, predicted protein sequences, functional annotation results, FASTA index files, a file manifest, and SHA256 checksums for haplotype 1 and haplotype 2. This record is intended to support the manuscript "Inter-haplotype inversions and repeat expansion in the sexually deceptive orchid Chiloglottis trapeziformis".
Stockner, Thomas · Al Makhlouf, Mounaf
1 files · 8.0 MB · gzip
This data record contains all input data and scripts to rerun the simulations and it includes the output structures. The dataset contains the raw and processed data used to create figures 5I, 5J of the assocated manuscript.
Bertola, Anouk · Wenner, Nicolas · LEMOS ROCHA, Leonardo Filipe · et al.
16 files · 8.0 MB · fastq, gzip, xlsx
Sequencing data and raw data for Bertola et al .
Tian, Lu · Yuxiu, Liu · Peng, zhao · et al.
2 files · 8.0 MB · gzip, xlsx
This dataset contains the imputed SNP genotypes of 406 bread wheat accessions in VCF format, used for haplotype-based GWAS analysis of root architecture.
Aalborg, Trine
7 files · 8.0 MB · csv, gzip
README - Data for "Genomic constraint and hypervariability in tetraploid potatoes" Trine Aalborg, May 2026 The data applied in the study includes phenotypic and genotypic information on the MASPOT panel (768 F1 progeny of an 18-parent diallel cross - property of Danespo A/S). The genotypic data was generated using genotyping-by-sequencing technology as described in the paper. Following genotype calling and filtration (5-60x read depth, < 50 % missing rate, > 1 % MAF), the total SNP set includes 151,164 biallelic SNPs. Coordinates of these SNPs relative to the DMv6.1 potato reference genome are provided. In addition to genotypes across the clones, the estimated GERP score of that SNP from (Wu et al., 2023) is reported. The manuscript analyses only consider markers with reliable GERP scores (MSA alignment depth > 50, and neutral score > 2), which corresponded to 97,815 of the total 151,164 biallelic SNPs. SNPeff annotations of the markers (based on the DMv6.1 reference genome) are also appended. The phenotypes were collected across 1-2 field trials, depending on the traits, and includes a minimum of two replicates per clone from a randomized block design. There are phenotypes for eight traits: dry matter content [%], yield (hkg/ha), senescence [1-9], flesh color [1-9], tubers/plant, tuber length [mm], tuber diameter [mm], and tuber size [mm^3]. Metadata includes phenotyping year and block location of the plot as well as pedigree of the diallel offspring. File descriptions: gt_MASPOT.csv - .csv file of the non-imputed genotypic data of 151,164 SNPs for the 768 F1 clones (those with GERP scores). Columns 1-3 are SNP coordinates and SNP IDs. Column names from column 4 and onwards are clone IDs. gt_MASPOT_imputed.csv - .csv file holding the imputed (random forest, missRanger algorithm) genotypic data of 151,164 SNPs for the 768 F1 clones. gt_MASPOT_recoded.csv - .csv file of the recoded, imputed genotypic data of 151,164 SNPs for the 768 F1 clones. The SNPs are recoded from original MASPOT ref/alt allele (based on AF in the MASPOT panel) to the alternative allele = the derived allele in the 100-Solanaceae panel. pt_MASPOT.csv - .csv file holding the phenotypic data (eight traits) of the 768 F1 clones (Clone_ID). The number of observations varies across traits. Also including metadata: year of phenotyping (Year), block number (Block, Line_in_block), clone parents (Mother, Father, Family). GERP_MASPOT.csv - .csv file holding the GERP scores (including alignment depth, neutral scores, and a marker annotation based on GERP score thresholds (deleterious, neutral, hypervariable, or low quality)), SNPeff annotations, and the derived allele in the 100-Solanaceae panel from (Wu et al., 2023) [MASPOT_Alt_Allele_Is_Sol_Derived_Allele - used for recoding of the genotypes] of the MASPOT SNPs with GERP scores. snps.MASPOT_F1.vcf.gz - zipped .vcf file of the GBS MASPOT genotypic data (both discrete genotype calls and allele frequencies) called to the DMv6.1 potato reference genome. Filtered to read depth 5x, MQ > 30. A total of 160,920 biallelic SNPs. Includes the 768 F1 progeny analyzed. Literature: Wu, Y., Li, D., Hu, Y., Li, H., Ramstein, G. P., Zhou, S., et al. (2023). Phylogenomic discovery of deleterious mutations facilitates hybrid potato breeding. Cell 186, 2313-2328.e15. doi: 10.1016/j.cell.2023.04.008
Salamzade, Rauf
4 files · 8.0 MB · gzip
Expected results from running comprehensive tests using zol v1.6.19. Script for running tests also included (but is primarily updated/tracked on zol's GitHub repo). The latest testing dataset is also included, correcting prior issues with corrupted FASTA files of Aspergillus genomes.
Jiménez-Blanco, Albert · López-Villellas, Lorién · Moure, Juan Carlos · et al.
1 files · 8.0 MB · gzip
Datasets used in the "Theseus: Fast and Optimal Affine-Gap Sequence-to-Graph Alignment" paper. The data is compressed into the theseus_datasets.tar.gz file. This file includes the datasets for the two experiments on the paper: MSA datasets: File Size Source hiv_100.fasta 100 sequences, each of approximately 10Kbp https://github.com/niemasd/ViralMSA/blob/master/example/example_hiv.fas mtb_benchmark_50kbp_shortened.fna 342 sequences, each of approximately 50Kbp Derived from a set of 342 RefSeq-complete whole-genome assemblies of Mycobacterium tuberculosis genomes. Sequences start at the dnaA gene and are 50Kbp long. mtb_benchmark_250kbp_trmB.fna 342 sequences, each of approximately 250Kbp Derived from a set of 342 RefSeq-complete whole-genome assemblies of Mycobacterium tuberculosis genomes. Sequences start at the dnaA gene and are truncated at the trmB gene. mtb_benchmark_500kbp_thiE.fna 342 sequences, each of approximately 500Kbp Derived from a set of 342 RefSeq-complete whole-genome assemblies of Mycobacterium tuberculosis genomes. Sequences start at the dnaA gene and are truncated at the thiE gene. mtb_benchmark_1Mbp_gltA2.fna 342 sequences, each of approximately 1Mbp Derived from a set of 342 RefSeq-complete whole-genome assemblies of Mycobacterium tuberculosis genomes. Sequences start at the dnaA gene and are truncated at the gltA2 gene. covid_19_complete.fasta 2732 sequences, each of approximately ∼30 Kbp in length 2732 GenBank-complete SARS-CoV-2 genome assemblies. monkeypox_100_seq.fasta 100 sequences, each of approximately ∼200 Kbp in length 100 RefSeq-complete whole genome assemblies of Monkey pox's virus. Sequence-to-graph /pangenome read mapping datasets: File Size Source SRR062634_1.filt_REDUCED.fasta 250K sequences of length 100bp Human Pangenome Reference Consortium 211109_M024_V350038332_L01_HUMuarfR092940-606_1_REDUCED.fasta 250K sequences of length 150bp NIST Genome in a Bottle (GIAB) project D1_S1_L001_R1_001_REDUCED.fasta 250K sequences of length 250bp NIST Genome in a Bottle (GIAB) project This dataset is derived from original data produced by third parties, as detailed above. All rights to the original data remain with the original authors or copyright holders. Users are responsible for ensuring compliance with the licensing terms of the original data sources. Sequence-to-cyclic-graph datasets: These datasets have been synthetically generated to test the ability of Theseus to align against cyclic reference graphs, comparing it to the graph unfolding strategy used by vg map. The synthetic cyclic graphs with a controlled structure. For that, we create backbone graphs with N in {10, 100, 1000} nodes, with node sequences of average length 50 bp, and connect each node to its successor to ensure graph connectivity. Then, we add a bounded number of random edges per node. These random edges can connect to nearby forward nodes or to previously created nodes, introducing cycles into the graph. For each graph size, we generate queries of length 100, 150, and 250 bp. We have nine pairs of reference graph and queries. This happens because we have three graph sizes, with N in {10, 100, 1000} nodes, and 3 query sizes, of 100, 150, and 250 base pairs. Code snapshot: This dataset also contains a snapshot of the code use to conduct the experiments in the manuscript, corresponding to Theseus v0.1 on Github.
Bhat, Shreeharsha G · Mahajan, Daanish · Jain, Chirag
7 files · 8.0 MB · gzip, zip
Panngenome graph files (GFA format) used for benchmarking panbubble and hairpin detection in the paper Billi: Provably Accurate and Scalable Bubble Detection in Pangenome Graphs . Includes minigraph pangenome graphs and pangene gene graphs. Note: The HPRC Minigraph-Cactus pangenome graphs (hprc-v1.1-mc-chm13 and hprc-v2.0-mc-chm13) and chromosome-level graphs (chrX and chrY) are not included due to their large size. The chromosome-level graphs (chrX, chrY) are available as .vg files and must be converted to GFA format using vg. Direct download links: hprc-v1.1-mc-chm13.gfa.gz: https://s3-us-west-2.amazonaws.com/human-pangenomics/pangenomes/freeze/freeze1/minigraph-cactus/hprc-v1.1-mc-chm13/hprc-v1.1-mc-chm13.gfa.gz hprc-v2.0-mc-chm13.gfa.gz: https://human-pangenomics.s3.amazonaws.com/pangenomes/scratch/2025_02_28_minigraph_cactus/hprc-v2.0-mc-chm13/hprc-v2.0-mc-chm13.gfa.gz chrX.vg: https://human-pangenomics.s3.amazonaws.com/pangenomes/scratch/2025_02_28_minigraph_cactus/hprc-v2.0-mc-chm13/hprc-v2.0-mc-chm13.chroms/chrX.vg chrY.vg: https://human-pangenomics.s3.amazonaws.com/pangenomes/scratch/2025_02_28_minigraph_cactus/hprc-v2.0-mc-chm13/hprc-v2.0-mc-chm13.chroms/chrY.vg Vg command to convert .vg to .gfa format: vg view input.vg > output.gfa
Koide, Kenji
2 files · 8.0 MB · gzip
A dataset for LiDAR-IMU-GNSS localization. This is a synthetic dataset generated with the Prius on Sonoma Raceway scene in Gazebo. All sequences are recoreded in the ROS2 bag format. Files rosbag2_2026_07_02-13_42_01 : Sequence 1 (209.6s) rosbag2_2026_07_02-13_51_33 : Sequence 2 (179.6s) sonoma.ply : Map data generated using GLIM running on Sequence 1 Topics /prius/cmd_vel | Type: geometry_msgs/msg/Twist /prius/gnss/fix | Type: sensor_msgs/msg/NavSatFix /prius/gnss/pose | Type: geometry_msgs/msg/PoseStamped /prius/imu | Type: sensor_msgs/msg/Imu /prius/odom | Type: nav_msgs/msg/Odometry /prius/points | Type: sensor_msgs/msg/PointCloud2 /tf | Type: tf2_msgs/msg/TFMessage /tf_static | Type: tf2_msgs/msg/TFMessage sonoma.ply
Shixian, Dong · Chengxin, Jiang · Caroline M., Eakin · et al.
4 files · 8.0 MB · gzip
This repository contains the ambient noise cross-correlation functions (CCFs) and dispersion curves (group and phase velocity) for the paper 'Characterizing Crustal Structure for Natural Hydrogen Exploration in the Southeastern Gawler Craton Using Adaptive Ambient Noise Tomography'.
Kovaliov, Michael
1 files · 100 MB · parquet