hybrid · semantic + lexical · 441 datasets ranked · 4.93s
Putri Lenggo Genni · Tavasya Alia Anjani · Muhammad Fasha Asshofa
3 files · 2.8 MB · csv, pdf, xlsx
Xiong, Mo · Xue, Ming
501 rows × 2 cols · 16 KB · csv
2 numeric
Drerup, Christian
8,000 rows × 9 cols · 511 KB · csv
5 numeric · 4 categorical
LI, Chuan
200 rows × 5 cols · 28 KB · csv
4 categorical · 1 text
Computational Network Science
1 files · 415 KB · gzip
The drug-target-interaction dataset is a combinatorial complex representing drug-target interactions and similarity relationships. The dataset is based on the work by Perlman et al., which combines multiple drug and gene similarity measures to predict drug-target interactions. Find the dataset details in AHORN .
Anonymous
1 files · 8.0 MB · gzip
This dataset contains 2134 unique malicious Python packages collected from PyPI, used for the evaluation of H2GLM (Hierarchical Heterogeneous Graph Learning framework enhanced by LLMs for Malicious package detection). Malicious samples are derived from a large-scale benchmark for malicious Python package detection, supplemented by packages collected from public security advisories and threat intelligence feeds. Deduplication was enforced via MD5 and fuzzy hashing. The archive malicious_2134.tar.gz contains the original .tar.gz distribution files as collected from PyPI. Intended for academic security research only.
Sasai, Kazuto · Yukio-Pegio, Gunji
7 files · 8.0 MB · gzip
# Data archive - Internal-state criticality in Bayesian-inverse-Bayesian inference **Paper.** *Internal-state criticality in Bayesian-inverse-Bayesian inference*, K. Sasai and Y.-P. Gunji (Physical Review Research, submitted). **Source repository.** <https://github.com/kazsasai/bayesian-inverse-bayesian-rps> **DOI.** `10.5281/zenodo.20533918`. --- ## Contents This deposit contains the raw simulation outputs underlying every data figure of the paper. Three archives are provided: two cover the main simulation data (*full reproducibility* vs *quick figure rebuild*), and one small archive holds the reinforcement-learning baseline-control data: | Archive | Size (compressed) | Contains | Use case | |---|---|---|---| | `paperA_data_full.tar.gz` | ~2.7 GB | Full simulation output tree (~16 GB uncompressed; 2713 files): all per-run JSONs and NPZs from `simulation/{reward_huge,nhand,reward_huge_v2,analyze_sharpness_plateau,reward}/data/` and `simulation_tie_mode_ablation/data/` | Independent re-analysis from raw outputs | | `paperA_data_figure_only.tar.gz` | ~1.1 GB | The 163 specific JSON/NPZ files actually read by `build_all.py` (~1.7 GB uncompressed) | Rebuild figures only | | `paperA_data_baseline_control.tar.gz` | ~21 MB | Pooled run-length arrays (`pnas_rl_comparison/data/baseline_dwells.npz`) for the RL-baseline control - WSLS, tabular Q-learning, and regret matching vs BIB; 40 seeds, T=2e5 - backing Fig. 4 (`fig_control_ab`) | Rebuild the RL-baseline control figure | The two main archives preserve the relative-path layout so that extracting either at `<repo>/data/` lets `build_all.py` find the data without further configuration. The baseline-control archive instead carries the `pnas_rl_comparison/data/...` path and extracts at the **repository root**. See **Reproducing the figures** below. Supporting files: * `MANIFEST_canonical.txt` - the in-repo data manifest (`BIB_Levy_v2/latex/figures/scripts/zenodo_data_manifest.txt`), listing each data tree, the figure(s) it feeds, and the generating script. * `figure_only_file_list.txt` - exhaustive 163-line list of relative paths inside `paperA_data_figure_only.tar.gz`, captured by auditing every `open()` call from a clean `build_all.py` run (and re-running with caches cleared so that no precomputed intermediate hid raw-data references). * `checksums.sha256` - SHA-256 of all three archives. ## Reproducing the figures Both tarballs preserve the same layout, so the workflow is identical: ```bash # 1. Clone the source repo git clone https://github.com/kazsasai/bayesian-inverse-bayesian-rps.git cd bayesian-inverse-bayesian-rps # 2. Get the data: pick ONE archive # (full = raw-output independent re-analysis; # figure-only = just enough to rebuild figures) mkdir -p data tar xzf /path/to/paperA_data_figure_only.tar.gz -C data # OR _full # 3. Install dependencies pip install numpy matplotlib powerlaw # 4. Rebuild figures python BIB_Levy_v2/latex/figures/scripts/build_all.py # (or run individual scripts: build_Fig3_universality.py, etc.) ``` Alternatively, point `PAPERA_DATA` at an extraction directory anywhere on disk: ```bash tar xzf paperA_data_figure_only.tar.gz -C /scratch/papera_data export PAPERA_DATA=/scratch/papera_data python BIB_Levy_v2/latex/figures/scripts/build_all.py ``` `figdata.py` in the source repo searches `$PAPERA_DATA`, then `<repo>/data/`, then the in-repo `simulation/` tree, in that order. ### RL-baseline control figure (Fig. 4) `paperA_data_baseline_control.tar.gz` carries the `pnas_rl_comparison/data/...` path, so extract it at the **repository root** (not `<repo>/data/`): ```bash tar xzf /path/to/paperA_data_baseline_control.tar.gz -C bayesian-inverse-bayesian-rps python pnas_si/figures/build_fig_control.py # -> fig_control_ab.{pdf,png} ``` The figure's BIB curves are read from the main data (the `reward_huge_*` `durations_bib-*` JSONs in the full / figure-only archive, via `$PAPERA_DATA`); the baseline curves come from the archive above. To regenerate the baseline data from scratch instead (deterministic, ~minutes): ```bash python pnas_rl_comparison/run_baseline_control.py # -> baseline_dwells.npz python pnas_rl_comparison/analyze_baseline_control.py ``` ## What `paperA_data_figure_only.tar.gz` excludes * The 17 G of per-run / per-step JSONs in the data trees that no current figure reads. * Intermediate caches (`fig*_ccdf_cache.json`) - these are regenerated by `build_Fig4_robustness.py` and `build_FigS2_nh_ccdf.py` on first run. * The small bundled inputs already shipped with the GitHub repo at `BIB_Levy_v2/latex/figures/scripts/data/` (`scheme_summary.csv`, `bo_tournament_results.json`, `sigma_*_rs_bib-bib.json`, `data_ivb_{equil,biased}.npz`). The build scripts read these straight from the repo. ## Verifying integrity ```bash shasum -a 256 -c checksums.sha256 ``` ## Citation If you use these data, please cite both the paper (forthcoming) and this Zenodo record. The repository's `README` is updated with the final citation on publication. ## License Data are released under CC-BY-4.0 (deposit metadata sets this on Zenodo). Source code in the GitHub repository is under its own LICENSE file.
guo, na · pan, jiachen · li, tiantian · et al.
14 files · 100 MB · csv, rar, tsv
ADM_LSIR is a physics-inspired laparoscopic aerosol degradation dataset for aerosol-aware surgical image analysis and image restoration. The v1.0.0 release contains: - 21,916 clean clinical laparoscopic frames (clean/) - 9,562 real intraoperative aerosol-degraded frames (degraded/) - 36,052 simulated aerosol masks, including 19,701 smoke-like masks and 16,351 trajectory masks (mask/) - Blender simulation/cache materials (ADM_LSIR_Blender_simulation_files_v1.0.rar) - metadata_quality_report_v1.0.csv - recommended_splits_v1.0.csv - video_mapping_v1.0.csv - parts_manifest.txt - checksums_v1.0.tsv - release_manifest_v1.0.json All released clinical frames are de-identified and stored as lossless PNG files. Filenames use anonymized video identifiers, e.g., C-V##-####.png for clean frames and D-V##-####.png for degraded frames. The recommended split is defined at the source_video_id/public_video_label level to reduce leakage across frames from the same source video. The public video labels in video_mapping_v1.0.csv provide privacy-safe source-video identifiers (video1-video19). The Blender archive documents the smoke and trajectory mask simulation setup and supports reuse, but it is not a guaranteed exact per-mask reproduction package. The released pre-rendered mask library is the primary reusable dataset component. Source code for synthesis and quality screening is available at: https://github.com/SweetDeathh/ADM_LSIR
Manghi, Paolo
2 files · 8.0 MB · gzip, tsv
Using gene-level and species-level shotgun metagenomics, we provide the first characterization of the rural, Bolivian microbiome; we identified microbial genes which strongly correlate (rho>0.45) with arsenic in urine, and that overall contribute to substantiate that the gut microbiome helps tolerate arsenic via a evict-out-of-house mechanism. Mediation analysis, followed by phylogenetic investigation of metagenomic-assembled genomes, further strengthens this observation. This study elucidates the role of the microbiome in helping to tolerate arsenic-rich environments, and paves the way for probiotic interventions that may mitigate the effects of this toxic metal. This repository contains 2,478 HQ MAGs from the MicroToxBol cohort in fasta format and a descriptive table comprising taxonomic annotation, quality-checks, and coverage estimation.
Goncalves dos Santos, Karen Cristine · Desgagné-Penix, Isabel · Merindol, Natacha
16 files · 8.0 MB · gzip
General description Compiled dataset of transcriptome assemblies, transcriptome annotation and expression quantification for 27 Amaryllidoideae species and 4 hybrid cultivars: Amaryllis belladonna , Clivia miniata , Crinum asiaticum , Crinum x powellii , Galanthus elwesii , Galanthus sp., Hippeastrum cv. Blossom Peacock, Hippeastrum cv. Jewel, Hippeastrum cv. Royal Velvet, Hippeastrum striatum , Hippeastrum vittatum , Leucojum aestivum , Lycoris aurea , Lycoris chinensis , Lycoris incarnata , Lycoris longituba , Lycoris radiata , Lycoris sprengeri , Narcissus aff pseudonarcissus , Narcissus papyraceus , Narcissus pseudonarcissus , Narcissus cv. Tête-à-Tête, Narcissus tazetta , Narcissus viridiflorus , Phycella aff cyrtanthoides , Rhodophiala pratensis , Scadoxus multiflorus , Traubia modesta , Zephyranthes candida , Zephyranthes carinata , Zephyranthes treatiae. Data includes transcriptome assemblies, Transdecoder predictions of peptide sequences (along with predicted coding seqeunces, GFF3 and BED files), expression quantification (count and TPM matrices for genes and trinity isoforms obtained with Kallisto and from Salmon), and annotation results (from EggNOG, Pfam, Uniprot Swissprot and Rfam), as well as signal peptide and transmemberane domain predictions (from SignalP and TmHMM). Annotations were compiled into a report for each species using Trinotate. For Lycoris aurea and Narcissus cv. Tête-à-Tête, there are two assemblies and corresponding files: Lycoris_aurea_PB and Lycoris_aurea_TH, and Narcissus_TaT_PB Narcissus_TaT_TH. PB assemblies were constructed solely with long-read sequencing data (PacBio, PB); while TH assemblies were constructed with short reads using Trinity, with long-read assembly being used for scaffolding step of Trinity (Trinity Hybrid, TH). Expression quantification was published on NCBI GEO (accessions GSE329951 , GSE329957 , GSE330014 and GSE331457 ). File descriptions All unitigs and predicted protein sequences are prefixed with an acronym to identify the species: Species Acronym NCBI TSA accession Amaryllis belladonna Ambel deposited on GSE331457 Clivia miniata Clmin DBNKRK000000000 Crinum asiaticum Crasi DBNIJL000000000 Crinum x powellii Crpow DBNKRS000000000 Galanthus elwesii Gaelw DBNKRO000000000 Galanthus sp. Gasp DBNKRR000000000 Hippeastrum cv. Blossom Peacock HispBP DBNMYE000000000 Hippeastrum cv. Jewel HispJW DBNMYC000000000 Hippeastrum cv. Royal Velvet HispRV DBNMYD000000000 Hippeastrum striatum Histr DBNIJG000000000 Hippeastrum vittatum Hivit DBNKRF000000000 Leucojum aestivum Leaes DBNIJJ000000000 Lycoris aurea (PB) Lyaur DBNFTY000000000 Lycoris aurea (TH) Lyaur DBNKRI000000000 Lycoris chinensis Lychi DBNKRE000000000 Lycoris incarnata Lyinc DBNKRD000000000 Lycoris longituba Lylon DBNKRG000000000 Lycoris radiata Lyrad deposited on GSE331457 Lycoris sprengeri Lyspr DBNKRL000000000 Narcissus aff pseudonarcissus Naafps DBNKRP000000000 Narcissus papyraceus Napap DBNKRN000000000 Narcissus pseudonarcissus Napse DBNIJK000000000 Narcissus Tête-à-Tête (PB) NaspPB DBNNFO000000000 Narcissus Tête-à-Tête (TH) NaspTH DBNUFN000000000 Narcissus tazetta Nataz DBNPME000000000 Narcissus viridiflorus Navir deposited on GSE331457 Phycella sp. Phsp deposited on GSE331457 Rhodophiala pratensis Rhpra deposited on GSE331457 Scadoxus multiflorus Scmul DBNIJH000000000 Traubia modesta Trmod deposited on GSE331457 Zephyranthes candida Zecan DBNKRH000000000 Zephyranthes carinata Zecar DBNIJI000000000 Zephyranthes treatiae Zetre deposited on GSE331457 Files Files are organized by type of analysis/data, meaning all expression quantification data obtained with Kallisto are compressed into a single file, all results from BlastP are in the same file, etc. Compressed file name Individual file type Content Blastp_Uniprot.tar.gz Tabular 33 files (1 per assembly) generated with BLASTP against Uniprot SwissProt release 2024_04. Files in blast output format 6 (standard columns). EggNOG_Emapper.tar.gz Tabular 33 tabular files (1 per assembly). Generated with eggnog.emapper (annotations file format described in eggnog.emapper's wiki ) Expression_Kallisto.tar.gz Tabular 128 files (4 per assembly, 2 for long-read assemblies): read counts and TPM values for both trinity "genes" and trinity "isoforms"). Row names are "gene"/ "isoform" IDs, column names are SRA run IDs. Expression_Salmon.tar.gz Tabular 128 files (4 per assembly, 2 for long-read assemblies): read counts and TPM values for both trinity "genes" and trinity "isoforms"). Row names are "gene"/ "isoform" IDs, column names are SRA run IDs. Final_assemblies.tar.gz Fasta 33 files (1 per assembly). Assemblies generated in this study, after contamination and expression filtering. HMMScan_PfamA.tar.gz Domain hits table 33 files (1 per assembly). Generated with hmmscan (option ‑‑domtblout, domain hits table, explained in hmmer's user guide [119] ) against the Pfam-A database. Infernal_Rfam.tar.gz Target hits table format 2 33 files (1 per assembly) generated with Infernal's cmscan (using the Trinotate wrapper) against the Rfam database. Table format 2 described in Infernal's user guide section 6 ). Metadata_studies.tar.gz Tabular (semi-colon separated columns) 31 tabular files (one per species/cultivar) indicating: Bioproject ID, SRA run ID, Sample name, Biosample ID, tissue, genotype (cultivar, when specified), treatment, batch, original publication citation, and DOI of original publication. Signalp6.tar.gz Tabular or GFF3 99 files (3 per assembly: prediction_results.txt, output.gff3 and region_output.gff3) generated with SignalP6. TmHMM2.tar.gz Tabular 33 files (1 per assembly) generated with TmHMM2 (short format, described in the guide tab of https://services.healthtech.dtu.dk/services/TMHMM-2.0/ Transcriptomes_unfiltered.tar.gz Fasta 33 files (1 per assembly). Unfiltered (prior to contamination and expression screening) assemblies generated in this study. Transdecoder_bed.tar.gz BED 33 BED files (1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam) Transdecoder_cds.tar.gz Fasta 33 fasta files (1 per assembly) of predicted coding sequences generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam) Transdecoder_gff3.tar.gz GFF3 33 GFF3 files (1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam). Transdecoder_proteomes.tar.gz Fasta 33 predicted proteome files (.pep, 1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam).
R. Panicker, Vyshakh · J. Smug, Bogna · Klein-Sousa, Victor · et al.
1 files · 8.0 MB · gzip
Data deposit for: Panicker VR, Smug BJ, Klein-Sousa V, Enright MC, Taylor NMI, Drulis-Kawa Z, Mostowy RJ. "Structural modularity of receptor-binding proteins underlies host-range strategy diversification in Klebsiella pneumoniae phages." bioRxiv 2026. https://doi.org/10.64898/2026.05.12.724579 This archive contains all large binary data files required to reproduce the analyses and figures in the paper above. The associated code is available at: https://github.com/VyshakhRP/RBP-div-hostrange Contents: - 01_raw/02_assemblies/ - Raw lysate assemblies for unpublished Klebsiella pneumoniae phages - 01_raw/03_host_genomes/ - Klebsiella pneumoniae host genome sequences (.fasta) - 01_raw/04_phage_genomes/ - Phage genome sequences (.fasta, .gb) - 02_intermediate/09_af3_predictions/ - AlphaFold 3 structure predictions for all receptor-binding proteins (.cif) - 02_intermediate/14_rbps-ecods/ecod.develop288.domains.txt - ECOD domain database (v20230309, develop288) - 02_intermediate/14_rbps-ecods/ecod.develop288.fasta.txt - ECOD domain FASTA (v20230309, develop288) - 05_output/01_genomes/ - Processed phage genome files - 05_output/03_proteins/ - Protein FASTAs and receptor-binding protein structures Usage: Download the archive, extract it, and place each subdirectory into the corresponding location in the cloned GitHub repository. Full instructions are provided in the repository README under Data Availability. Note: The ECOD database files can alternatively be downloaded directly from http://prodata.swmed.edu/ecod/distributions/ (v20230309 / develop288).
Requena Rolanía, Jose María · Greif, Gonzalo · ROBELLO, CARLOS
1 files · 8.0 MB · fasta
This dataset contains the genome sequence for Trypanosoma cruzi (strain Dm28c). This genome sequence was de novo assembled using PacBio Hi-Fi and Illumina sequencing platforms by Greif et al (2026. PMID: 41501640). The genome was assembled into 32 contigs, which represent complete chromosomes. The provided Fasta file also contains an additional contig corresponding to the maxicircle (mitochondrial genome) sequence. The Fasta files included in this dataset were downloaded from GenBank (assembly GCA_044048535.1; May 22, 2026). Additional information about the Dm28cT2T genome assembly and gene annotations may be accessed through the link: https://cruzi.pasteur.uy/
Bar, Ido
9 files · 8.0 MB · gff, gzip
These files are downsampled WGS fastq files (250k paired-end reads each) of a fungal pathogen ( Ascochyta rabiei ), generated on MGI DNSeq-T7. The files are intended to be used directly as test datasets in PopFun - a Nextflow pipeline for variant calling in fungal genomes.
Zhang, Zirui · Jones, Ashley · Schwessinger, Benjamin · et al.
1 files · 8.0 MB · gzip
This dataset contains annotation supporting files for the chromosome-scale, haplotype-resolved genome assembly of the sexually deceptive Australian orchid Chiloglottis trapeziformis. The archive includes haplotype-specific BRAKER gene annotation files, predicted coding sequences, predicted protein sequences, functional annotation results, FASTA index files, a file manifest, and SHA256 checksums for haplotype 1 and haplotype 2. This record is intended to support the manuscript "Inter-haplotype inversions and repeat expansion in the sexually deceptive orchid Chiloglottis trapeziformis".
Burman, Nathaniel · Buyukyoruk, Murat · Wiegand, Tanner · et al.
4 files · 8.0 MB · fasta
This folder contains a multiple sequence alignment of Cas7 homologs in .fasta format, the domain-level annotations from PFAM and CasFinder, and an associated phylogenetic tree in .newick format.
Stockner, Thomas · Al Makhlouf, Mounaf
1 files · 8.0 MB · gzip
This data record contains all input data and scripts to rerun the simulations and it includes the output structures. The dataset contains the raw and processed data used to create figures 5I, 5J of the assocated manuscript.
Winter, Henry · Severino, MaryKay · Volunteer Scientist
16 files · 100 MB · csv, pdf, zip
These are audio recordings taken by an Eclipse Soundscapes (ES) Data Collector during the week of the April 08, 2024 Total Solar Eclipse. It was decided to include only raw, unprocessed audio data files in each site-specific ZIP archive and in each Zenodo record. This decision was so that any researcher can independently verify, reproduce, and extend the analysis performed. As a result, some sites have WAV files with 0 bytes of data or timestamps outside the range of probable recording times. Procedures used by the Eclipse Soundscapes team to process audio data for its purposes are outlined in the Data Management reports located in the Eclipse Soundscapes Zenodo community. Data with 0 bytes of data were included for completeness. Data Site location information: Latitude: 44.46311 Longitude: -71.68203 Local Eclipse Type: Total Solar Eclipse Solar Eclipse Eclipse Percent (%): 100 WAV files Time & Date Settings: Set with Automated AudioMoth Time Chime (More information on TimeStamp Setting below) Data Collector Start Time Notes: N/A Included Data: Audio files in WAV format with the date and time in UTC within the file name: YYYYMMDD_HHMMSS meaning YearMonthDay_HourMinuteSecond For example, 20240411_141600.WAV means that this audio file starts on April 11, 2024 at 14:16:00 Coordinated Universal Time (UTC) CONFIG Text file: Includes AudioMoth device setting information, such as sample rate in Hertz (Hz), gain, firmware, etc. README.md: Markdown formatted file with information about the recording and recording site. file_list.csv: A machine and human file that gives the following information on each file in the record: File Name, File Type, Description, File Size in kilobytes, Name of Associated Data Dictionary with the file, calculated SHA-512 Hash of the file as a unique identifier to insure data integrity during transfer and compression. total_eclipse_data.csv: A machine and human readable file that gives the following information about the site where the audio data recording was taken: ESID#, Latitude, Longitude, Eclipse_type, CoveragePercent, Eclipse Start UTC (1st contact), Totality Start UTC (2nd contact), Totality End UTC (3rd Contact), Eclipse End UTC (4th Contact), Max Eclipse Time UTC License.txt: A human readable file that explains the terms and conditions under which the data can be used. AudioMoth_Operation_Manual.pdf: A human readable document that explains the use of an AudioMoth device. The document is current up to the time of the AudioMoth's use in the Eclipse Soundscapes project. file_list_data_dict.csv: A machine and human data dictionary file that gives information on the variables contained within the file_list.csv file. CONFIG_data_dict.csv: A machine and human data dictionary file that gives information on the variables contained within the CONFIG.TXT file. eclipse_data_data_dict.csv: A machine and human data dictionary file that gives information on the variables contained within the total_eclipse_data.csv file. WAV_data_dict.csv: A machine and human data dictionary file that gives information on the variables contained within the *.WAV files. ES_Data_Management_Pre-Eclipse_Data_Infrastructure_Stage_0.pdf: PDF document that describes Stage 0 (Pre-Eclipse Infrastructure and Data Stewardship Planning) of the Eclipse Soundscapes (ES) data lifecycle. ES_Data_Management_Receipt_Sorting_and_Metadata_Organization_Stage_1.pdf: PDF document that describes Stage 1 (Receipt, Sorting, and Metadata Organization) of the Eclipse Soundscapes (ES) data lifecycle. ES_Data_Management_Data_Processing_Stage_2.pdf: PDF document that describes Stage 2 (Data Processing) of the Eclipse Soundscapes (ES) data Volunteer Scientists. 2023 and 2024 solar eclipse soundscapes audio datalifecycle. ES_Data_Management_Data_Sharing_Stage_3.pdf: PDF document that describes Stage 3 (Public Data Sharing) of the Eclipse Soundscapes (ES) data lifecycle. Eclipse Information for this location: Eclipse Date: April 08, 2024 Eclipse Start Time (UTC) (1st Contact): 18:16:15 Totality Start Time (UTC) (2nd Contact): [N/A if partial eclipse] 19:28:56 Eclipse Maximum Time [when the most possible amount of the Sun in blocked] (UTC): 19:29:25 Totality End Time (UTC) (3rd Contact): [N/A if partial eclipse] 19:29:53 Eclipse End Time (UTC) (4th Contact): [N/A if partial eclipse] 20:38:30 Audio Data Collection During Eclipse Week ES Data Collectors used AudioMoth devices to record audio data, known as soundscapes, over a 5-day period during the eclipse week: 2 days before the eclipse, the day of the eclipse, and 2 days after. The complete raw audio data collected by the Data Collector at the location mentioned above is provided here. This data may or may not cover the entire requested timeframe due to factors such as availability, technical issues, or other unforeseen circumstances. ES ID# Information: Each AudioMoth recording device was assigned a unique Eclipse Soundscapes Identification Number (ES ID#). This identifier connects the audio data, submitted via a MicroSD card, with the latitude and longitude information provided by the data collector through an online form. The ES team used the ES ID# to link the audio data with its corresponding location information and then uploaded this raw audio data and location details to Zenodo. This process ensures the anonymity of the ES Data Collectors while allowing them to easily search for and access their audio data on Zenodo. TimeStamp Information: The ES team and the Data Collectors took care to set the date and time on the AudioMoth recording devices using an AudioMoth time chime before deployment, ensuring that the recordings would have an automatic timestamp. However, participants also manually noted the date and start time as a backup in case the time chime setup failed. The notes above indicate whether the WAV audio files for this site were timestamped manually or with the automated AudioMoth time chime. Common Timestamp Error: Some AudioMoth devices experienced a malfunction where the timestamp on audio files reverted to a date in 1970 or before, even after initially recording correctly. Despite this issue, the affected data was still included in this ES site's collected raw audio dataset. Latitude & Longitude Information: The latitude and longitude for each site was taken manually by data collectors and submitted to the ES team, either via a web form or on paper. It is shared in Decimal Degrees format. General Project Information: The Eclipse Soundscapes Project is a NASA Volunteer Science project funded by NASA Science Activation that is studying how eclipses affect life on Earth during the October 14, 2023 annular solar eclipse and the April 8, 2024 total solar eclipse. Eclipse Soundscapes revisits an eclipse study from almost 100 years ago that showed that animals and insects are affected by solar eclipses! Like this study from 100 years ago, ES asked for the public's help. ES uses modern technology to continue to study how solar eclipses affect life on Earth! Eclipse Soundscapes is an enterprise of ARISA Lab, LLC and is supported by NASA award No. 80NSSC21M0008. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Aeronautics and Space Administration. Eclipse map/figure/table/predictions courtesy of Fred Espenak, NASA/Goddard Space Flight Center, from eclipse.gsfc.nasa.gov . Eclipse Data Version Definitions {1st digit = year, 2nd digit = Eclipse type (1=Total Solar Eclipse, 9=Annular Solar Eclipse, 0=Partial Solar Eclipse), 3rd digit is unused and in place for future use} 2023.9.0 = Week of October 14, 2023 Annular Eclipse Audio Data, Path of Annularity (Annular Eclipse) 2023.0.0 = Week of October 14, 2023 Annular Eclipse Audio Data, OFF the Path of Annularity (Partial Eclipse) 2024.1.0 = Week of April 8, 2024 Total Solar Eclipse Audio Data, Path of Totality (Total Solar Eclipse) 2024.0.0 = Week of April 8, 2024 Total Solar Eclipse Audio Data , OFF the Path of Totality (Partial Solar Eclipse) *Please note that this dataset's version number is listed below. Eclipse Soundscapes Data Collector Role Training and Implementation Resources Manual (2023-2024) (Archival Copy) This site-level record includes the Eclipse Soundscapes Data Collector Role Training and Implementation Resources Manual (2023-2024) . The manual documents the participant training, device setup procedures, metadata submission requirements, ES ID system, timestamp protocols, data return workflow, and public archiving processes used during the October 14, 2023 annular solar eclipse and the April 8, 2024 total solar eclipse. The manual is preserved for transparency and reproducibility and reflects the procedures under which this dataset was collected and processed. (DOI 10.5281/zenodo.18623442) Data Receipt, Processing, and Analysis Methods All programs supporting Stages 1–3 are openly available in the: Eclipse Soundscapes GitHub repository: https://github.com/ARISA-Lab-LLC/ESCSP Data Management Lifecycle The following section documents the relationship of this record to the full Eclipse Soundscapes (ES) data lifecycle, a multi-stage workflow designed to support large-scale participatory science, long-term data stewardship, open science, and scientific reuse. Each stage addressed a different operational need, beginning before eclipse deployment and continuing through validation, public archiving, and scientific analysis. Together, these stages transformed distributed volunteer-submitted audio recordings into structured, documented, publicly accessible NASA-funded research assets. Stage 0: Pre-Eclipse Infrastructure and Deployment Preparation Severino, M., & Winter, H. (2026). Eclipse Soundscapes Data Management: Pre-Eclipse Infrastructure and Deployment Preparation (Stage 0). Zenodo. https://doi.org/10.5281/zenodo.20413370 Stage 0 focused on building the operational foundation required to support geographically distributed eclipse data collection at national scale. This stage included AudioMoth device preparation, accessibility modifications, ES ID # assignment systems, metadata collection workflows, participant training materials, deployment logistics, and planning for downstream data stewardship and archival workflows. The 2023 annular eclipse served as both a scientific investigation and a large-scale operational beta test that informed improvements for the 2024 total solar eclipse campaign. Related Citations and Resources: Severino, M., & Kline, T. (2025, November 24). Eclipse Soundscapes Apprentice Role Curriculum: Solar Eclipses and Multi-Sensory Observing (Informal Education). Zenodo. https://doi.org/10.5281/zenodo.17703003 Severino, M., & Bauer, D. J. (2026). Eclipse Soundscapes Observer Role Training and Resources Manual (2023–2024). Zenodo. https://doi.org/10.5281/zenodo.18633602< /li> Severino, M., Winter, H., & Bauer, D. J. (2026). Eclipse Soundscapes Data Collector Role Training and Implementation Manual (2023–2024). Zenodo. https://doi.org/10.5281/zenodo.18623443 Stage 1: Receipt, Sorting, and Metadata Organization Severino, M., & Winter, H. (2026). Eclipse Soundscapes Data Management: Receipt, Sorting, and Metadata Organization (Stage 1). Zenodo. https://doi.org/10.5281/zenodo.19471425 Stage 1 transformed returned participant materials into organized, traceable site-level records. This included receiving mailed microSD cards, consolidating participant-submitted metadata, reconciling handwritten and online records, organizing physical audio media by ES ID #, and deriving eclipse timing and coverage information using NASA eclipse prediction datasets. The outputs of Stage 1 established the structured metadata relationships required for downstream validation, processing, archiving, and analysis workflows. Related Citations and Resources: Winter, H., & Goncalves, J. (2026). EPTT (Eclipse Phase Timing Tool) [Computer software]. GitHub. https://github.com/ARISA-Lab-LLC/ESCSP-Eclipse-Phase-Timing-Tool /li> Espenak, F. (n.d.). Eclipse predictions by Fred Espenak, NASA's GSFC Eclipse Web Site. NASA Goddard Space Flight Center. http://eclipse.gsfc.nasa.gov/eclipse.html Stage 2: Data Processing and Validation Severino, M., & Winter, H. (2026). Eclipse Soundscapes Data Management: Data Processing (Stage 2). Zenodo. https://doi.org/10.5281/zenodo.18683402 Stage 2 focused on centralized audio ingestion, validation, timestamp verification, metadata reconciliation, and preparation of datasets for analysis and public sharing. During this stage, returned audio recordings were processed using custom open-source tools developed by the ES team, including ES WAVES and ES AMES. The project implemented scalable infrastructure capable of processing large volumes of participant-submitted microSD cards while preserving all raw audio data without modification. Stage 2 established the validated dataset structure required for long-term preservation and scientific analysis. Related Citations and Resources: Winter, H., & Goncalves, J. (2026). ES WAVES (Eclipse Soundscapes WAV Audio Validation & Extraction Suite) [Computer software]. GitHub. https://github.com/ARISA-Lab-LLC/ESCSP-ES-WAV-Audio-Validation-Extraction-Suite Winter, H., & Goncalves, J. (2026). ES AMES (Eclipse Soundscapes AudioMoth Metadata Extractor Suite) [Computer software]. GitHub. https://github.com/ARISA-Lab-LLC/ESCSP-ES-AMES-AudioMoth-Metadata-Extractor-Suite Stage 3: Public Data Sharing and Open Archiving Severino, M., & Winter, H. (2026). Eclipse Soundscapes Data Management: Public Audio Data Sharing (Stage 3). Zenodo. https://doi.org/10.5281/zenodo.18683437 Stage 3 transformed validated site-level datasets into publicly archived, DOI-assigned research records published through the Eclipse Soundscapes Zenodo Community. This stage included dataset packaging, metadata standardization, README generation, integrity verification, DOI assignment, and automated repository upload workflows using the Automated Zenodo Upload Software (AZUS). These workflows established the project's long-term open-science infrastructure and ensured that datasets remained findable, accessible, interoperable, reusable, and citable for future scientific and educational use. Related Citations and Resources: Winter, H., & Goncalves, J. (2026). AZUS (Automated Zenodo Upload Software) [Computer software]. GitHub. https://github.com/ARISA-Lab-LLC/AZUS-Automated-Zenodo-Upload-Software Stage 4: Scientific Analysis and Research Use Stage 4 involves the scientific analysis and interpretation of validated eclipse soundscape datasets. Analysis workflows utilized datasets verified during earlier stages to investigate eclipse-related environmental and animal vocalization changes across hundreds of recording sites. This stage also includes broader scientific interpretation, publication development, and continued reuse of Eclipse Soundscapes datasets and infrastructure for future research, education, and open-science applications. Related Citations and Resources: Pease, B., Gilbert, N., & Severino, M. (2026). Eclipse Soundscapes Preliminary Findings – How Eclipses Affect Nature as determined by Sound (Recorded Webinar). Zenodo. https://doi.org/10.5281/zenodo.18613979 Gilbert, N. A., Pease, B. S., Severino, M., & Winter, H. III. (2026). Photic niche explains avian behavioral responses to solar eclipses. Ecology and Evolution, 16(2), e73090. https://doi.org/10.1002/ece3.73090 Analysis code repository: https://github.com/BrentPease1/eclipse-traits Companion Zenodo record archiving structured analysis scripts and derived outputs: https://doi.org/10.5281/zenodo.15790879[r] Public Archiving, Privacy, and Data Transparency The Eclipse Soundscapes Data Collector Role Training and Implementation Manual (2023-2024) includes a detailed explanation of how Eclipse Soundscapes audio data are publicly archived on Zenodo, how participant privacy is protected through the ES ID system, and how transparency and traceability are maintained. It also outlines the criteria for determining which recordings are included in the public archive, as well as the distinction between publicly shared archival data and datasets used for ES-led scientific analyses. Participants and data users can consult this section for full documentation of the project's open science and privacy practices. Severino, M., & Winter, H. (2026). Eclipse Soundscapes Data Collector Role Training and Implementation Manual (2023–2024). Zenodo. https://doi.org/10.5281/zenodo.18623443 Citations Individual Site Citation: APA Citation (7th edition) Winter, H., Severino, M., & Volunteer Scientist. (2026). 2024 solar eclipse soundscapes audio data [Audio dataset, ES ID# 292]. Zenodo.{Insert DOI} Collected by volunteer scientists as part of the Eclipse Soundscapes Project. This project is supported by NASA award No. 80NSSC21M0008. Eclipse Community Citation Winter, H., Severino, M., & Volunteer Scientists. 2023 and 2024 solar eclipse soundscapes audio data [Collection of audio datasets]. Eclipse Soundscapes Community, Zenodo. https://zenodo.org/communities/eclipsesoundscapes/ Collected by volunteer scientists as part of the Eclipse Soundscapes Project This project is supported by NASA award No. 80NSSC21M0008.
Fischer, Fabian Jörg · Chave, Jerome · Zanne, Amy · et al.
45 cols · 8.0 MB · csv
31 categorical · 12 numeric · 2 text
The Global Wood Density Database v.2 The Global Wood Density Database v.2 (GWDD v.2) is a collection of 109,626 taxonomically standardized wood density records and 15,093 additional bark density records. Data include measurements at different levels of aggregation (individual, species) and both georeferenced records and values from the literature. For a full description of the database, please see the corresponding manuscript. ( Fischer et al. 2026. Beyond species means - the intraspecific contribution to global wood density variation. New Phytol. https://doi.org/10.1111/nph.70860 ). It includes and supersedes the GWDD v.1 ( Zanne et al. 2009 , https://doi.org/10.5061/dryad.234 ). When using the GWDD v.2 in your work, please cite Fischer et al. 2026 as well as this repository using the corresponding DOI (10.5281/zenodo.16919509). If you would like to report an issue or suggest improvements for future updates of the GWDD, please do so on github: https://github.com/fischer-fjd/GWDD/issues Aggregated wood density data We provide pre-aggregated wood density data, with wood density estimates at species, binomial species and genus level. Wood density estimates are derived from hierarchical (random effects) models and provided both as simple species mean values ( wsg_est ) and as species mean values for trunks ( wsg_est_trunk ) and branches ( wsg_est_branch ) separately. In addition, we provide raw wood density means ( wsg_raw ), but we do not recommend using them for practical purposes due to outliers for poorly sampled species. gwddagg_v2.x_species: pre-aggregated wood density values for 17,261 species, including infraspecific epithets; comprises 16,828 taxonomically resolved species well as 433 values with uncertain taxonomic status gwddagg_v2.x_binomial: pre-aggregated wood density values for 16,905 binomial species gwddagg_v2.x_genus: pre-aggregated wood density values for 3,198 genera Raw wood density data In addition, we also provide the underlying raw wood density database. This collection contains one metadata file and raw data files in .csv format. Since special characters (e.g., in the references) can be distorted by operating systems when reading in .csv files, we also provide all data in .RData format, which can be loaded into R with the load() function. columns_gwdd_v2.x : metadata for all the columns included in the GWDD v.2 gwdd_v2.x : the GWDD v.2, including all 109,626 wood density records gwdd_v2.x_withbark : the GWDD v.2, including all 109,626 wood density records and 15,093 additional bark density records
Bertola, Anouk · Wenner, Nicolas · LEMOS ROCHA, Leonardo Filipe · et al.
16 files · 8.0 MB · fastq, gzip, xlsx
Sequencing data and raw data for Bertola et al .
Tian, Lu · Yuxiu, Liu · Peng, zhao · et al.
2 files · 8.0 MB · gzip, xlsx
This dataset contains the imputed SNP genotypes of 406 bread wheat accessions in VCF format, used for haplotype-based GWAS analysis of root architecture.