LI, Chuan
200 rows × 5 cols · 28 KB · csv
4 categorical · 1 text
hybrid · semantic + lexical · 323 datasets ranked · 0.85s
LI, Chuan
200 rows × 5 cols · 28 KB · csv
4 categorical · 1 text
guo, na · pan, jiachen · li, tiantian · et al.
14 files · 100 MB · csv, rar, tsv
ADM_LSIR is a physics-inspired laparoscopic aerosol degradation dataset for aerosol-aware surgical image analysis and image restoration. The v1.0.0 release contains: - 21,916 clean clinical laparoscopic frames (clean/) - 9,562 real intraoperative aerosol-degraded frames (degraded/) - 36,052 simulated aerosol masks, including 19,701 smoke-like masks and 16,351 trajectory masks (mask/) - Blender simulation/cache materials (ADM_LSIR_Blender_simulation_files_v1.0.rar) - metadata_quality_report_v1.0.csv - recommended_splits_v1.0.csv - video_mapping_v1.0.csv - parts_manifest.txt - checksums_v1.0.tsv - release_manifest_v1.0.json All released clinical frames are de-identified and stored as lossless PNG files. Filenames use anonymized video identifiers, e.g., C-V##-####.png for clean frames and D-V##-####.png for degraded frames. The recommended split is defined at the source_video_id/public_video_label level to reduce leakage across frames from the same source video. The public video labels in video_mapping_v1.0.csv provide privacy-safe source-video identifiers (video1-video19). The Blender archive documents the smoke and trajectory mask simulation setup and supports reuse, but it is not a guaranteed exact per-mask reproduction package. The released pre-rendered mask library is the primary reusable dataset component. Source code for synthesis and quality screening is available at: https://github.com/SweetDeathh/ADM_LSIR
Winter, Henry · Severino, MaryKay · Volunteer Scientist
16 files · 100 MB · csv, pdf, zip
These are audio recordings taken by an Eclipse Soundscapes (ES) Data Collector during the week of the April 08, 2024 Total Solar Eclipse. It was decided to include only raw, unprocessed audio data files in each site-specific ZIP archive and in each Zenodo record. This decision was so that any researcher can independently verify, reproduce, and extend the analysis performed. As a result, some sites have WAV files with 0 bytes of data or timestamps outside the range of probable recording times. Procedures used by the Eclipse Soundscapes team to process audio data for its purposes are outlined in the Data Management reports located in the Eclipse Soundscapes Zenodo community. Data with 0 bytes of data were included for completeness. Data Site location information: Latitude: 44.46311 Longitude: -71.68203 Local Eclipse Type: Total Solar Eclipse Solar Eclipse Eclipse Percent (%): 100 WAV files Time & Date Settings: Set with Automated AudioMoth Time Chime (More information on TimeStamp Setting below) Data Collector Start Time Notes: N/A Included Data: Audio files in WAV format with the date and time in UTC within the file name: YYYYMMDD_HHMMSS meaning YearMonthDay_HourMinuteSecond For example, 20240411_141600.WAV means that this audio file starts on April 11, 2024 at 14:16:00 Coordinated Universal Time (UTC) CONFIG Text file: Includes AudioMoth device setting information, such as sample rate in Hertz (Hz), gain, firmware, etc. README.md: Markdown formatted file with information about the recording and recording site. file_list.csv: A machine and human file that gives the following information on each file in the record: File Name, File Type, Description, File Size in kilobytes, Name of Associated Data Dictionary with the file, calculated SHA-512 Hash of the file as a unique identifier to insure data integrity during transfer and compression. total_eclipse_data.csv: A machine and human readable file that gives the following information about the site where the audio data recording was taken: ESID#, Latitude, Longitude, Eclipse_type, CoveragePercent, Eclipse Start UTC (1st contact), Totality Start UTC (2nd contact), Totality End UTC (3rd Contact), Eclipse End UTC (4th Contact), Max Eclipse Time UTC License.txt: A human readable file that explains the terms and conditions under which the data can be used. AudioMoth_Operation_Manual.pdf: A human readable document that explains the use of an AudioMoth device. The document is current up to the time of the AudioMoth's use in the Eclipse Soundscapes project. file_list_data_dict.csv: A machine and human data dictionary file that gives information on the variables contained within the file_list.csv file. CONFIG_data_dict.csv: A machine and human data dictionary file that gives information on the variables contained within the CONFIG.TXT file. eclipse_data_data_dict.csv: A machine and human data dictionary file that gives information on the variables contained within the total_eclipse_data.csv file. WAV_data_dict.csv: A machine and human data dictionary file that gives information on the variables contained within the *.WAV files. ES_Data_Management_Pre-Eclipse_Data_Infrastructure_Stage_0.pdf: PDF document that describes Stage 0 (Pre-Eclipse Infrastructure and Data Stewardship Planning) of the Eclipse Soundscapes (ES) data lifecycle. ES_Data_Management_Receipt_Sorting_and_Metadata_Organization_Stage_1.pdf: PDF document that describes Stage 1 (Receipt, Sorting, and Metadata Organization) of the Eclipse Soundscapes (ES) data lifecycle. ES_Data_Management_Data_Processing_Stage_2.pdf: PDF document that describes Stage 2 (Data Processing) of the Eclipse Soundscapes (ES) data Volunteer Scientists. 2023 and 2024 solar eclipse soundscapes audio datalifecycle. ES_Data_Management_Data_Sharing_Stage_3.pdf: PDF document that describes Stage 3 (Public Data Sharing) of the Eclipse Soundscapes (ES) data lifecycle. Eclipse Information for this location: Eclipse Date: April 08, 2024 Eclipse Start Time (UTC) (1st Contact): 18:16:15 Totality Start Time (UTC) (2nd Contact): [N/A if partial eclipse] 19:28:56 Eclipse Maximum Time [when the most possible amount of the Sun in blocked] (UTC): 19:29:25 Totality End Time (UTC) (3rd Contact): [N/A if partial eclipse] 19:29:53 Eclipse End Time (UTC) (4th Contact): [N/A if partial eclipse] 20:38:30 Audio Data Collection During Eclipse Week ES Data Collectors used AudioMoth devices to record audio data, known as soundscapes, over a 5-day period during the eclipse week: 2 days before the eclipse, the day of the eclipse, and 2 days after. The complete raw audio data collected by the Data Collector at the location mentioned above is provided here. This data may or may not cover the entire requested timeframe due to factors such as availability, technical issues, or other unforeseen circumstances. ES ID# Information: Each AudioMoth recording device was assigned a unique Eclipse Soundscapes Identification Number (ES ID#). This identifier connects the audio data, submitted via a MicroSD card, with the latitude and longitude information provided by the data collector through an online form. The ES team used the ES ID# to link the audio data with its corresponding location information and then uploaded this raw audio data and location details to Zenodo. This process ensures the anonymity of the ES Data Collectors while allowing them to easily search for and access their audio data on Zenodo. TimeStamp Information: The ES team and the Data Collectors took care to set the date and time on the AudioMoth recording devices using an AudioMoth time chime before deployment, ensuring that the recordings would have an automatic timestamp. However, participants also manually noted the date and start time as a backup in case the time chime setup failed. The notes above indicate whether the WAV audio files for this site were timestamped manually or with the automated AudioMoth time chime. Common Timestamp Error: Some AudioMoth devices experienced a malfunction where the timestamp on audio files reverted to a date in 1970 or before, even after initially recording correctly. Despite this issue, the affected data was still included in this ES site's collected raw audio dataset. Latitude & Longitude Information: The latitude and longitude for each site was taken manually by data collectors and submitted to the ES team, either via a web form or on paper. It is shared in Decimal Degrees format. General Project Information: The Eclipse Soundscapes Project is a NASA Volunteer Science project funded by NASA Science Activation that is studying how eclipses affect life on Earth during the October 14, 2023 annular solar eclipse and the April 8, 2024 total solar eclipse. Eclipse Soundscapes revisits an eclipse study from almost 100 years ago that showed that animals and insects are affected by solar eclipses! Like this study from 100 years ago, ES asked for the public's help. ES uses modern technology to continue to study how solar eclipses affect life on Earth! Eclipse Soundscapes is an enterprise of ARISA Lab, LLC and is supported by NASA award No. 80NSSC21M0008. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Aeronautics and Space Administration. Eclipse map/figure/table/predictions courtesy of Fred Espenak, NASA/Goddard Space Flight Center, from eclipse.gsfc.nasa.gov . Eclipse Data Version Definitions {1st digit = year, 2nd digit = Eclipse type (1=Total Solar Eclipse, 9=Annular Solar Eclipse, 0=Partial Solar Eclipse), 3rd digit is unused and in place for future use} 2023.9.0 = Week of October 14, 2023 Annular Eclipse Audio Data, Path of Annularity (Annular Eclipse) 2023.0.0 = Week of October 14, 2023 Annular Eclipse Audio Data, OFF the Path of Annularity (Partial Eclipse) 2024.1.0 = Week of April 8, 2024 Total Solar Eclipse Audio Data, Path of Totality (Total Solar Eclipse) 2024.0.0 = Week of April 8, 2024 Total Solar Eclipse Audio Data , OFF the Path of Totality (Partial Solar Eclipse) *Please note that this dataset's version number is listed below. Eclipse Soundscapes Data Collector Role Training and Implementation Resources Manual (2023-2024) (Archival Copy) This site-level record includes the Eclipse Soundscapes Data Collector Role Training and Implementation Resources Manual (2023-2024) . The manual documents the participant training, device setup procedures, metadata submission requirements, ES ID system, timestamp protocols, data return workflow, and public archiving processes used during the October 14, 2023 annular solar eclipse and the April 8, 2024 total solar eclipse. The manual is preserved for transparency and reproducibility and reflects the procedures under which this dataset was collected and processed. (DOI 10.5281/zenodo.18623442) Data Receipt, Processing, and Analysis Methods All programs supporting Stages 1–3 are openly available in the: Eclipse Soundscapes GitHub repository: https://github.com/ARISA-Lab-LLC/ESCSP Data Management Lifecycle The following section documents the relationship of this record to the full Eclipse Soundscapes (ES) data lifecycle, a multi-stage workflow designed to support large-scale participatory science, long-term data stewardship, open science, and scientific reuse. Each stage addressed a different operational need, beginning before eclipse deployment and continuing through validation, public archiving, and scientific analysis. Together, these stages transformed distributed volunteer-submitted audio recordings into structured, documented, publicly accessible NASA-funded research assets. Stage 0: Pre-Eclipse Infrastructure and Deployment Preparation Severino, M., & Winter, H. (2026). Eclipse Soundscapes Data Management: Pre-Eclipse Infrastructure and Deployment Preparation (Stage 0). Zenodo. https://doi.org/10.5281/zenodo.20413370 Stage 0 focused on building the operational foundation required to support geographically distributed eclipse data collection at national scale. This stage included AudioMoth device preparation, accessibility modifications, ES ID # assignment systems, metadata collection workflows, participant training materials, deployment logistics, and planning for downstream data stewardship and archival workflows. The 2023 annular eclipse served as both a scientific investigation and a large-scale operational beta test that informed improvements for the 2024 total solar eclipse campaign. Related Citations and Resources: Severino, M., & Kline, T. (2025, November 24). Eclipse Soundscapes Apprentice Role Curriculum: Solar Eclipses and Multi-Sensory Observing (Informal Education). Zenodo. https://doi.org/10.5281/zenodo.17703003 Severino, M., & Bauer, D. J. (2026). Eclipse Soundscapes Observer Role Training and Resources Manual (2023–2024). Zenodo. https://doi.org/10.5281/zenodo.18633602< /li> Severino, M., Winter, H., & Bauer, D. J. (2026). Eclipse Soundscapes Data Collector Role Training and Implementation Manual (2023–2024). Zenodo. https://doi.org/10.5281/zenodo.18623443 Stage 1: Receipt, Sorting, and Metadata Organization Severino, M., & Winter, H. (2026). Eclipse Soundscapes Data Management: Receipt, Sorting, and Metadata Organization (Stage 1). Zenodo. https://doi.org/10.5281/zenodo.19471425 Stage 1 transformed returned participant materials into organized, traceable site-level records. This included receiving mailed microSD cards, consolidating participant-submitted metadata, reconciling handwritten and online records, organizing physical audio media by ES ID #, and deriving eclipse timing and coverage information using NASA eclipse prediction datasets. The outputs of Stage 1 established the structured metadata relationships required for downstream validation, processing, archiving, and analysis workflows. Related Citations and Resources: Winter, H., & Goncalves, J. (2026). EPTT (Eclipse Phase Timing Tool) [Computer software]. GitHub. https://github.com/ARISA-Lab-LLC/ESCSP-Eclipse-Phase-Timing-Tool /li> Espenak, F. (n.d.). Eclipse predictions by Fred Espenak, NASA's GSFC Eclipse Web Site. NASA Goddard Space Flight Center. http://eclipse.gsfc.nasa.gov/eclipse.html Stage 2: Data Processing and Validation Severino, M., & Winter, H. (2026). Eclipse Soundscapes Data Management: Data Processing (Stage 2). Zenodo. https://doi.org/10.5281/zenodo.18683402 Stage 2 focused on centralized audio ingestion, validation, timestamp verification, metadata reconciliation, and preparation of datasets for analysis and public sharing. During this stage, returned audio recordings were processed using custom open-source tools developed by the ES team, including ES WAVES and ES AMES. The project implemented scalable infrastructure capable of processing large volumes of participant-submitted microSD cards while preserving all raw audio data without modification. Stage 2 established the validated dataset structure required for long-term preservation and scientific analysis. Related Citations and Resources: Winter, H., & Goncalves, J. (2026). ES WAVES (Eclipse Soundscapes WAV Audio Validation & Extraction Suite) [Computer software]. GitHub. https://github.com/ARISA-Lab-LLC/ESCSP-ES-WAV-Audio-Validation-Extraction-Suite Winter, H., & Goncalves, J. (2026). ES AMES (Eclipse Soundscapes AudioMoth Metadata Extractor Suite) [Computer software]. GitHub. https://github.com/ARISA-Lab-LLC/ESCSP-ES-AMES-AudioMoth-Metadata-Extractor-Suite Stage 3: Public Data Sharing and Open Archiving Severino, M., & Winter, H. (2026). Eclipse Soundscapes Data Management: Public Audio Data Sharing (Stage 3). Zenodo. https://doi.org/10.5281/zenodo.18683437 Stage 3 transformed validated site-level datasets into publicly archived, DOI-assigned research records published through the Eclipse Soundscapes Zenodo Community. This stage included dataset packaging, metadata standardization, README generation, integrity verification, DOI assignment, and automated repository upload workflows using the Automated Zenodo Upload Software (AZUS). These workflows established the project's long-term open-science infrastructure and ensured that datasets remained findable, accessible, interoperable, reusable, and citable for future scientific and educational use. Related Citations and Resources: Winter, H., & Goncalves, J. (2026). AZUS (Automated Zenodo Upload Software) [Computer software]. GitHub. https://github.com/ARISA-Lab-LLC/AZUS-Automated-Zenodo-Upload-Software Stage 4: Scientific Analysis and Research Use Stage 4 involves the scientific analysis and interpretation of validated eclipse soundscape datasets. Analysis workflows utilized datasets verified during earlier stages to investigate eclipse-related environmental and animal vocalization changes across hundreds of recording sites. This stage also includes broader scientific interpretation, publication development, and continued reuse of Eclipse Soundscapes datasets and infrastructure for future research, education, and open-science applications. Related Citations and Resources: Pease, B., Gilbert, N., & Severino, M. (2026). Eclipse Soundscapes Preliminary Findings – How Eclipses Affect Nature as determined by Sound (Recorded Webinar). Zenodo. https://doi.org/10.5281/zenodo.18613979 Gilbert, N. A., Pease, B. S., Severino, M., & Winter, H. III. (2026). Photic niche explains avian behavioral responses to solar eclipses. Ecology and Evolution, 16(2), e73090. https://doi.org/10.1002/ece3.73090 Analysis code repository: https://github.com/BrentPease1/eclipse-traits Companion Zenodo record archiving structured analysis scripts and derived outputs: https://doi.org/10.5281/zenodo.15790879[r] Public Archiving, Privacy, and Data Transparency The Eclipse Soundscapes Data Collector Role Training and Implementation Manual (2023-2024) includes a detailed explanation of how Eclipse Soundscapes audio data are publicly archived on Zenodo, how participant privacy is protected through the ES ID system, and how transparency and traceability are maintained. It also outlines the criteria for determining which recordings are included in the public archive, as well as the distinction between publicly shared archival data and datasets used for ES-led scientific analyses. Participants and data users can consult this section for full documentation of the project's open science and privacy practices. Severino, M., & Winter, H. (2026). Eclipse Soundscapes Data Collector Role Training and Implementation Manual (2023–2024). Zenodo. https://doi.org/10.5281/zenodo.18623443 Citations Individual Site Citation: APA Citation (7th edition) Winter, H., Severino, M., & Volunteer Scientist. (2026). 2024 solar eclipse soundscapes audio data [Audio dataset, ES ID# 292]. Zenodo.{Insert DOI} Collected by volunteer scientists as part of the Eclipse Soundscapes Project. This project is supported by NASA award No. 80NSSC21M0008. Eclipse Community Citation Winter, H., Severino, M., & Volunteer Scientists. 2023 and 2024 solar eclipse soundscapes audio data [Collection of audio datasets]. Eclipse Soundscapes Community, Zenodo. https://zenodo.org/communities/eclipsesoundscapes/ Collected by volunteer scientists as part of the Eclipse Soundscapes Project This project is supported by NASA award No. 80NSSC21M0008.
Fischer, Fabian Jörg · Chave, Jerome · Zanne, Amy · et al.
45 cols · 8.0 MB · csv
31 categorical · 12 numeric · 2 text
The Global Wood Density Database v.2 The Global Wood Density Database v.2 (GWDD v.2) is a collection of 109,626 taxonomically standardized wood density records and 15,093 additional bark density records. Data include measurements at different levels of aggregation (individual, species) and both georeferenced records and values from the literature. For a full description of the database, please see the corresponding manuscript. ( Fischer et al. 2026. Beyond species means - the intraspecific contribution to global wood density variation. New Phytol. https://doi.org/10.1111/nph.70860 ). It includes and supersedes the GWDD v.1 ( Zanne et al. 2009 , https://doi.org/10.5061/dryad.234 ). When using the GWDD v.2 in your work, please cite Fischer et al. 2026 as well as this repository using the corresponding DOI (10.5281/zenodo.16919509). If you would like to report an issue or suggest improvements for future updates of the GWDD, please do so on github: https://github.com/fischer-fjd/GWDD/issues Aggregated wood density data We provide pre-aggregated wood density data, with wood density estimates at species, binomial species and genus level. Wood density estimates are derived from hierarchical (random effects) models and provided both as simple species mean values ( wsg_est ) and as species mean values for trunks ( wsg_est_trunk ) and branches ( wsg_est_branch ) separately. In addition, we provide raw wood density means ( wsg_raw ), but we do not recommend using them for practical purposes due to outliers for poorly sampled species. gwddagg_v2.x_species: pre-aggregated wood density values for 17,261 species, including infraspecific epithets; comprises 16,828 taxonomically resolved species well as 433 values with uncertain taxonomic status gwddagg_v2.x_binomial: pre-aggregated wood density values for 16,905 binomial species gwddagg_v2.x_genus: pre-aggregated wood density values for 3,198 genera Raw wood density data In addition, we also provide the underlying raw wood density database. This collection contains one metadata file and raw data files in .csv format. Since special characters (e.g., in the references) can be distorted by operating systems when reading in .csv files, we also provide all data in .RData format, which can be loaded into R with the load() function. columns_gwdd_v2.x : metadata for all the columns included in the GWDD v.2 gwdd_v2.x : the GWDD v.2, including all 109,626 wood density records gwdd_v2.x_withbark : the GWDD v.2, including all 109,626 wood density records and 15,093 additional bark density records
Aalborg, Trine
7 files · 8.0 MB · csv, gzip
README - Data for "Genomic constraint and hypervariability in tetraploid potatoes" Trine Aalborg, May 2026 The data applied in the study includes phenotypic and genotypic information on the MASPOT panel (768 F1 progeny of an 18-parent diallel cross - property of Danespo A/S). The genotypic data was generated using genotyping-by-sequencing technology as described in the paper. Following genotype calling and filtration (5-60x read depth, < 50 % missing rate, > 1 % MAF), the total SNP set includes 151,164 biallelic SNPs. Coordinates of these SNPs relative to the DMv6.1 potato reference genome are provided. In addition to genotypes across the clones, the estimated GERP score of that SNP from (Wu et al., 2023) is reported. The manuscript analyses only consider markers with reliable GERP scores (MSA alignment depth > 50, and neutral score > 2), which corresponded to 97,815 of the total 151,164 biallelic SNPs. SNPeff annotations of the markers (based on the DMv6.1 reference genome) are also appended. The phenotypes were collected across 1-2 field trials, depending on the traits, and includes a minimum of two replicates per clone from a randomized block design. There are phenotypes for eight traits: dry matter content [%], yield (hkg/ha), senescence [1-9], flesh color [1-9], tubers/plant, tuber length [mm], tuber diameter [mm], and tuber size [mm^3]. Metadata includes phenotyping year and block location of the plot as well as pedigree of the diallel offspring. File descriptions: gt_MASPOT.csv - .csv file of the non-imputed genotypic data of 151,164 SNPs for the 768 F1 clones (those with GERP scores). Columns 1-3 are SNP coordinates and SNP IDs. Column names from column 4 and onwards are clone IDs. gt_MASPOT_imputed.csv - .csv file holding the imputed (random forest, missRanger algorithm) genotypic data of 151,164 SNPs for the 768 F1 clones. gt_MASPOT_recoded.csv - .csv file of the recoded, imputed genotypic data of 151,164 SNPs for the 768 F1 clones. The SNPs are recoded from original MASPOT ref/alt allele (based on AF in the MASPOT panel) to the alternative allele = the derived allele in the 100-Solanaceae panel. pt_MASPOT.csv - .csv file holding the phenotypic data (eight traits) of the 768 F1 clones (Clone_ID). The number of observations varies across traits. Also including metadata: year of phenotyping (Year), block number (Block, Line_in_block), clone parents (Mother, Father, Family). GERP_MASPOT.csv - .csv file holding the GERP scores (including alignment depth, neutral scores, and a marker annotation based on GERP score thresholds (deleterious, neutral, hypervariable, or low quality)), SNPeff annotations, and the derived allele in the 100-Solanaceae panel from (Wu et al., 2023) [MASPOT_Alt_Allele_Is_Sol_Derived_Allele - used for recoding of the genotypes] of the MASPOT SNPs with GERP scores. snps.MASPOT_F1.vcf.gz - zipped .vcf file of the GBS MASPOT genotypic data (both discrete genotype calls and allele frequencies) called to the DMv6.1 potato reference genome. Filtered to read depth 5x, MQ > 30. A total of 160,920 biallelic SNPs. Includes the 768 F1 progeny analyzed. Literature: Wu, Y., Li, D., Hu, Y., Li, H., Ramstein, G. P., Zhou, S., et al. (2023). Phylogenomic discovery of deleterious mutations facilitates hybrid potato breeding. Cell 186, 2313-2328.e15. doi: 10.1016/j.cell.2023.04.008
Kovaliov, Michael
1 files · 100 MB · parquet
Omukuti, Rodney
1 files · 8.0 MB · vcf
Tsui, Claire · Briga, Michael · Komdeur, Jan · et al.
56 files · 8.9 MB · csvdeclared
Data and code for analysis in manuscript titled "Asynchrony of ageing among traits in a wild bird population" Dataframes ending with"_28_5.csv" and "survival_model.csv" are used in scripts model1-13, of which the output is plotted using "new model outputs.R" Code for Figures 1 and 2 are in script "new model outputs.R" asymmetry bivar ver3.R runs the bivariate models used to estimate the degree of synchrony of ageing. scripts starting with "aic.." are used in the analysis for age by lifespan interaction
Madhusudan, Gujral
6 files · 29 MB · parquetdeclared
Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.
Krajnc, Matej · Comi, Troy · Miao, Siqi · et al.
10 files · 25 GB · csv, zipdeclared
This dataset accompanies the manuscript "A Controlled in Silico Benchmark for GNN Prediction of Tissue Dynamics." It contains model prediction outputs, trained checkpoints, train/validation/test splits, spring-embedding outputs, generated analysis figures, analysis tables, and manuscript-specific diagnostic outputs used to reproduce the post-prediction analyses and figures. The dataset is distributed as logical ZIP archives with file-level and archive-level SHA-256 checksums. For questions, contact Tomer Stern at tomers@umich.edu.
MENEGUZZO, FRANCESCO · Zabini, Federica
4 files · 146 KB · csvdeclared
This record contains the de-identified, analysis-ready dataset and Python analysis code associated with a school-based quasi-experimental pilot study on therapist-guided forest therapy and persistent anxiety symptoms in adolescents. The study involved two fourth-year high-school classrooms in Cecina, Tuscany, Italy: one intervention classroom that attended four therapist-guided forest therapy sessions in a coastal pine forest, and one control classroom that followed usual school activities. The dataset includes anonymized student codes, classroom allocation, SCAS total raw scores and derived reduction scores across repeated assessments, POMS-A acute mood-state variables for the intervention classroom, exploratory post-intervention nature-exposure and connectedness variables, and exploratory school-performance variables. The repository includes four files: 1. forest_therapy_adolescent_anxiety_dataset.csv: de-identified analysis-ready dataset. 2. README_forest_therapy_adolescent_anxiety_dataset.txt: dataset metadata and variable descriptions. 3. analyze_forest_therapy_adolescent_anxiety.py: Python script used to reproduce the main analyses and generate analysis outputs. 4. README_analyze_forest_therapy_adolescent_anxiety.txt: operational guide for running the analysis script and interpreting the output files. The dataset uses semicolon-separated values. Missing values are encoded as NaN and must not be interpreted as zero. Because the study involved minors, item-level questionnaire responses and more granular school records are not shared; only de-identified and analysis-ready variables compatible with privacy, ethical approval, and consent constraints are provided.
Soltes, Julian
3 files · 12 MB · csv, pdfdeclared
This paper presents a non-parametric topography of the cosmological landscape, following the principle of maximum entropy to map the relative probability of all isotropic spacetime configurations. The topography is defined as an 11-D probability distribution, derived from a maximally unbiased ensemble of solutions to the Einstein Field Equations (EFE). Uncertainty prevents infinite precision, blurring the ensemble to quantify the relative likelihood of any physical microstate. The result provides a statistical foundation for the manifestation of our universe and its contents, described by probabilities and their respective gradients. This assumes unconstrained numerical coverage of the EFE landscape, mapping mathematically valid solutions that may violate parametric energy conditions. Uncertainty distorts the boundaries between these exotic states and the canonical phase, rationalizing their existence within our universe as quantifiable improbabilities. Resources: The computational implementation and main data ensemble are attached here, as well as maintained at: https://github.com/jgsoltes/Universe-KDE Researchers are encouraged to apply the QUASAR optimizer to their own high-dimensional or non-convex function landscapes. Source code and documentation are maintained at: https://github.com/jgsoltes/hdim-opt
Hackl, Jürgen
6 files · 132 MB · parquet, tiffdeclared
EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).
Fobbe, Sean
16 files · 1.4 GB · csv, pdf, zipdeclared
Überblick Das Corpus des deutschen Bundesrechts (C-DBR) ist eine möglichst vollständige Sammlung der konsolidierten Fassungen aller Gesetze und Verordnungen auf Bundesebene. Der Datensatz nutzt als seine Datenquelle das amtliche Internetangebot www.gesetze-im-internet.de des Bundesministeriums der Justiz und wertet dieses vollständig aus. Bitte lesen Sie zuerst das beiliegende Codebook! Es enthält wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante für Sie am besten geeignet ist. In der Regel empfehle ich für quantitative Forschung die CSV-Dateien und für traditionelle Forschung die PDF-Sammlung. Um das Gesetzgebungsverfahren näher zu beleuchten können Sie zusätzlich auf folgende Datensätze zurückgreifen (jeweils mit Links auf vergleichbare Datensätze anderer Autor:innen): Corpus der Drucksachen des Deutschen Bundestages (CDRS-BT) Corpus der Plenarprotokolle des Deutschen Bundestages (CPP-BT) Aktualisierung Dieser Datensatz wird ca. alle 3 Monate aktualisiert. Benachrichtigungen über neue und aktualisierte Datensätze veröffentliche ich immer zeitnah auf Mastodon unter @seanfobbe@fediscience.org NEU in Version 2026-07-09 Vollständige Aktualisierung der Daten Ausführung der Pipeline in Userland im Container Neues Skript für Zenodo Upload Eckdaten Stichtag: 9. Juli 2026 Umfang: 6.124 Bundesgesetze und -verordnungen der Bundesrepublik Deutschland Formate: CSV, PDF, EPUB, TXT und XML Features Einfache Nutzung für statistische Analysen mit CSV-Dateien Bis zu 42 Variablen in den CSV-Varianten Fortlaufende Aktualisierung Urheberrechtsfreiheit Sowohl für traditionelle Rechtsanwender als auch für Legal Tech-Anwendungen geeignete Formate (CSV, PDF, EPUB, TXT und XML) Umfangreicher Compilation Report um den Erstellungs-Prozess zu erläutern Hochauflösende Diagramme und deskriptive Tabellen für alle Zwecke Diagramme in PDF (Druck) und PNG (Web) verfügbar, Tabellen als menschen- und maschinenlesbares CSV Vollständiges tabellarisches Verzeichnis aller Rechtsakte und der vom BMJV gebrauchten Abkürzungen Netzwerk-Strukturen für alle Rechtsakte und Visualisierungen für über 1000 Rechtsakte (experimentell) Veröffentlichung des Source Codes Source Code und Compilation Report Der gesamte Erstellungs-Prozess ist vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollständigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (ähnlich dem Codebook). Zudem werden Robustness Checks auf Vollständigkeit und Plausibilität durchgeführt und in einem separaten Bericht dokumentiert. Der Compilation Report enthält den Code für die vollständige Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Er ist zusammen mit dem Source Code hinterlegt. Wenn Sie sich für Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst. Der vollständige Source Code - sowohl für die Erstellung des Datensatzes, als auch für das Codebook - ist öffentlich einsehbar und dauerhaft erreichbar im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt: https://zenodo.org/doi/10.5281/zenodo.4072934 Kryptographische Signaturen Die Integrität und Echtheit der einzelnen Archive des Datensatzes sind durch eine Zwei-Phasen-Signatur sichergestellt. In Phase I werden während der Kompilierung für jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert. In Phase II werden diese CSV-Datei und der Compilation Report mit meinem persönlichen geheimen GPG-Schlüssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgeführt werden kann, insbesondere im Rahmen von Replikationen, die persönliche Gewähr für Ergebnisse aber dennoch vorhanden ist. Die während der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Prüfsummen ist mit meiner persönlichen GPG-Signatur versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten: Name: Sean Fobbe (fobbe-data@posteo.de) Fingerabdruck: FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42 Kein Urheberrecht: Public Domain An den Normtexten und Metadaten besteht gem. § 5 Abs. 1 UrhG kein Urheberrecht, da sie amtliche Werke sind. § 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "Sächsischer Ausschreibungsdienst"). Alle eigenen Beiträge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gemäß einer CC0 1.0 Universal Public Domain License vollständig urheberrechtsfrei. Disclaimer Dieser Datensatz ist eine private wissenschaftliche Initiative und steht in keiner Verbindung zu Behörden, Gerichten oder anderen öffentlichen Stellen der Bundesrepublik Deutschland. Alternativen [Ab 10.06.2019, nur XML] Beckedorf, Janis/Coupette, Corinna/Hartung, Dirk. 2020. "gesetze-im-internet: A daily archive of https://www.gesetze-im-internet.de". GitHub. https://github.com/QuantLaw/gesetze-im-internet [Änderungsgesetze] Wehrmeyer, Stefan/Semsrott, Arne/Filter, Johannes. 2021. "OffeneGesetze.de ist eine zivilgesellschaftliche, ehrenamtliche Plattform für amtliche Gesetzesblätter". Open Knowledge Foundation. https://offenegesetze.de/ [Alte Rechtsakte] Open Knowledge Foundation. 2013. "Bundesgit". GitHub. https://github.com/bundestag/gesetze Weitere Open Access Veröffentlichungen (Fobbe) Website - www.seanfobbe.de Open Data - zenodo.org/communities/sean-fobbe-data/ Source Code - zenodo.org/communities/sean-fobbe-code/ Volltexte regulärer Publikationen - zenodo.org/communities/sean-fobbe-publications/ Kontakt Fehler gefunden? Anregungen? Kommentieren Sie gerne im Issue Tracker oder kontaktieren Sie mich über www.seanfobbe.de
Vera, Lourdes
31 files · 4.6 MB · csv, geojson, pngdeclared
Data, analysis scripts, and derived outputs reproducing the setback analysis between occupied buildings and active oil and gas wells in Karnes County, Texas, using public data (Texas Railroad Commission well locations; FEMA/ORNL USA Structures building footprints; and U.S. Census TIGER/Line block groups and 2020-2024 American Community Survey). The pipeline runs offline in Python (geopandas) and reproduces every reported distance, summary statistic, table, and figure in the associated article. This package reproduces and updates Chapter 3 of the author's doctoral dissertation: Vera, Lourdes (2022), "Environmental Data Justice in Action: Civically Valid Air Monitoring Near Oil and Gas Extraction in the Eagle Ford Shale Play," Ph.D. dissertation, Northeastern University, Boston, MA.
Wareesri, Prapassorn · Ieamvijarn, Subunn
14 files · 1.7 MB · csvdeclared
Dataset and replication code for the paper "The Regime-Dependent Value of Macroeconomic Information in Gold Futures Volatility Forecasting: A HAR-Machine Learning Comparison on COMEX" submitted to Investment Management and Financial Innovations. Data sourced from Yahoo Finance covering January 2014 to December 2025.
Haghbin, Kourosh
7 files · 121 KB · csv, xlsxdeclared
Supplementary code and data for the Master's thesis "The Influence of Green Infrastructure on the Urban Acoustic Environment: A Seasonal Noise Analysis in Bochum" (TU Dortmund University, 2026). The repository contains Python scripts implementing the full statistical analysis pipeline: site-level data assembly, Ordinary Least Squares (OLS) regression, and Multiscale Geographically Weighted Regression (MGWR) of BirdNET-derived normalised Shannon bird diversity (BN_H) against acoustic and structural green infrastructure predictors across 118 spring morning monitoring sites in Bochum, Germany. The primary script ( run_ols_mgwr_dawn.py ) implements the four-predictor model (dB, NDSI, AEI, and GI Tier; OLS R² = 0.319, MGWR R² = 0.796), selected as the primary model based on AICc (ΔAICc = -105.70 over the five-predictor alternative). A comparison script ( run_ols_mgwr_area.py ) implements the five-predictor model including log-transformed GI polygon area (OLS R² = 0.333, MGWR R² = 0.895). All analyses are implemented from scratch in Python 3 using NumPy and Pandas, without proprietary GIS or statistical libraries. The site-level analytical dataset (Dawn_4to9_Analysis_Data.xlsx, n = 118 sites, 04:00-09:00 dawn chorus window) is included to enable full reproduction of reported results.
Hagen, Karl · Zieher, Thomas · Stary, Ulrike · et al.
21 files · 6.6 MB · csv, pdfdeclared
Groundwater and landslide displacement records of the Eggerberg slope (Gradenbach landslide, Carinthia, Austria) Austrian Research Centre for Forests (BFW) Institute for Natural Hazards, Unit of Torrent Process & Hydrology K. Hagen, T. Zieher, U. Stary‚ E. Lang, S. Riedl, G. Priesch, J. Pichler, J. Rojacher First published June 2026 Contact: wasser.naturgefahren@bfw.gv.at Overview This dataset documents long-term hydrogeological and displacement monitoring at the deep-seated rock slide Eggerberg, situated in the Gradenbach catchment (Carinthia, Austria). The active mass movement covers approximately 2 km² and reaches depths exceeding 130 m. Owing to its interaction with the Gradenbach torrent system, the instability represents a significant natural hazard for nearby settlements in the Möll Valley. The dataset includes groundwater level and temperature measurements, as well as landslide displacement records collected by the Austrian Research Centre for Forests (BFW). It represents one of the longest continuous hydrogeological and geotechnical monitoring programs of a deep-seated gravitational slope deformation in the European Alps. It extends the existing dataset published in Hormes et al. (2026) by data of several boreholes, not used in the publication. Available groundwater temperature records were compiled and added in the borehole records. Furthermore, the displacement record of the extensometer was reset to zero after the data gap from 1995 to 1999. Dataset Description Groundwater monitoring Groundwater levels were monitored in 15 boreholes including several paired installations designed to observe groundwater conditions in different depth horizons and aquifer systems. Monitoring of the goundwater levels started in 1979, using manual cable light-plummet measurements with an accuracy of approximately 1 cm, and ended in 2024. In the beginning, measurements were generally performed every two weeks, since 1998 usually weekly. Groundwater temperature was measured between 1998 and 2015 with a sensor accuracy of 0.1°C. Continuous digital monitoring of groundwater level and temperature is additionally available for selected boreholes from July 2007 to December 2024 using OTT Orpheus Mini pressure probes. The observation series reveal the presence of approximately four hydrogeologically distinct aquifers within the moving rock mass. Landslide displacement monitoring Landslide displacement was monitored using a wire extensometer installed across the Gradenbach ravine. The instrument measured changes in the distance between the moving landslide mass and the comparatively stable opposite valley flank. The observation record extends from May 1979 to April 2024 and includes the same principal data gap between 1996 and 1998. Initially, measurements were recorded using analogue strip-chart systems and later digitized. In 2006, the monitoring station was upgraded with a continuous digital acquisition system (Thalimedes). To reduce short-term thermal effects caused by steel-wire expansion and contraction, the published dataset contains daily aggregated displacement values. All records (groundwater level and temperature, landslide displacement) were generally quality-controlled and checked for outliers and plausibility before publication. However, no warranty is given regarding their accuracy, completeness, or fitness for any particular purpose. Data Structure Groundwater Files Each borehole is provided as an individual CSV file (GRD-GWL-[borehole number]): • Column 1: Date (YYYY-MM-DD) • Column 2: Groundwater temperature (°C) • Column 3: Groundwater level below ground surface (m) Missing or unreliable values are coded as 9999. Groundwater levels are reported as negative values relative to the ground surface. Extensometer File The landslide displacement dataset is provided as a single CSV file (GRD_EXT.csv): • Column 1: Date (YYYY-MM-DD) • Column 2: Landslide displacement (cm), expressed as reduction of the distance between canyon slopes • Column 3: Measurement and data-quality code The PDF document 'GRD_Zenodo-V1-20260701.pdf' provides site information, methods, data gaps, and corrected observations. The ancillary document 'monitoring_data_gradenbach.html' provides minimal code snippets for working with the data in R. It further includes interactive plots of the data for visual inspection. Scientific R elevance The Gradenbach-Eggerberg dataset provides a unique long-term record enables the investigations of groundwater-controlled slope acceleration processes, and temporal trends in aquifer behavior. The dataset therefore constitutes an important resource for landslide process research, hazard assessment, hydrogeological investigations, and model validation.
Robert Koch-Institut
9 files · 54 MB · csv, pdf, zipdeclared
Der Datensatz "Intensivkapazitäten und COVID-19-Intensivbettenbelegung in Deutschland" des Robert Koch-Instituts dokumentiert die tägliche intensivmedizinische Versorgungslage seit der COVID-19-Pandemie. Basierend auf Meldungen aller intensivbettenführenden Krankenhäuser in Deutschland erfasst das DIVI-Intensivregister Echtzeitdaten zu belegten und freien Intensivbetten. Die Erhebung differenziert nach Altersgruppen, Regionen und Versorgungsstufen. COVID-19-Fälle auf Intensivstationen werden gesondert ausgewiesen. Die Daten stehen aggregiert auf Bundes-, Landes- und Kreisebene zur Verfügung. Damit bildet der Datensatz eine Grundlage für die Überwachung von Kapazitäten, die Koordination von Behandlungskapazitäten und politische Entscheidungsprozesse während der Pandemie und darüber hinaus.
Pifferi, Gabriele · Thulinsson, Felix · Söderlund, Niclas · et al.
9 files · 165 KB · csvdeclared
# Camera Monitor System (CMS) Dataset and Analysis Scripts Version 4 Dataset and analysis scripts associated with the manuscript: "Effects of Camera Height, Field of View, and Driver Age on Depth Judgement and Lane-Change Decisions in Camera Monitor Systems" ## Changelog ### Version 4 - Added the R analysis scripts for the analyses reported in the associated manuscript (`CMS_continuous_age_analysis.R`, `Within_subject_figures.R`, `confidence_analysis.R`, `unsigned_error_analysis.R`); see the Analysis scripts section below. - Data files unchanged from Version 3. ### Version 3 - Corrected a file export error that caused row truncation in `distance_estimation_cleaned.csv` and `lane_change_dataset.csv` (earlier versions). Most data columns were missing from each row in those files due to a comma/semicolon delimiter collision during export. Files have been rebuilt from the original source data. - Corrected a mixed decimal notation issue (comma vs dot) in the `Confidence` column of `distance_estimation_cleaned.csv`, which caused the column to be read as a string rather than a numeric type in standard CSV parsers. - Updated `variable_dictionary.csv` to accurately reflect the column names used in all three data files (earlier versions listed intended names that did not match the actual file contents). - Added two participants (P48, P49) to `participant_metadata.csv` whose records were missing from the earlier export due to a data extraction error. Their experimental task data were present in `lane_change_dataset.csv` but lacked corresponding metadata rows. - Clarified that the N=56 figure in the preprocessing section applies specifically to the distance estimation task (task-specific outlier exclusion). The lane-change task and participant metadata retain the full N=58. ### Version 2 Same datafiles as in Version was uploaded by mistake ### Version 1 Initial release. --- ## Licence This dataset and the accompanying analysis scripts are shared under the Creative Commons Attribution 4.0 International (CC BY 4.0) licence. https://creativecommons.org/licenses/by/4.0/ --- ## Included files ### Data #### participant_metadata.csv Participant-level demographic and background information. N=58. #### distance_estimation_cleaned.csv Cleaned participant-level data from the distance estimation task. N=56 after task-specific outlier exclusion (see Data preprocessing below). #### lane_change_dataset.csv Participant-level data from the lane-change task. N=58. #### variable_dictionary.csv Definitions and descriptions of all variables included across the three data files. ### Analysis scripts (R) #### CMS_continuous_age_analysis.R #### Within_subject_figures.R #### confidence_analysis.R #### unsigned_error_analysis.R See the Analysis scripts section below for descriptions, requirements, and run order. --- ## Experimental overview A controlled laboratory experiment combining two aligned data collections was conducted using dynamic rearward driving scenarios in a Camera Monitor System (CMS) environment. Participants completed: - a distance estimation task - a lane-change task (last safe gap, LSG) Experimental factors: - Field of View (FOV): 40°, 76°, 112° - Camera height: High / Low - Driver age: continuous variable (range 22-64 years, mean 38.2 years). The variable `AgeGroupMedianSplit` is retained in the data files for continuity with prior analyses (Pifferi, 2025), but is not used in the primary analysis reported in the associated manuscript. --- ## Data preprocessing Preliminary analyses examined the distribution of the dependent variables for the distance estimation (Dist) and lane-change (LSG) tasks. The LSG data did not show substantial skewness or kurtosis, whereas the Dist data contained several extreme values reflecting substantial underestimation relative to the actual target distances. Outliers in the distance estimation data were identified using a threshold of ±2.5 standard deviations from the mean of the dependent variables (absolute and relative distance estimation errors). Exclusion limits: - Absolute distance error: -84 m to 79 m - Relative distance error: -2.76 to 2.55 Two participants were excluded from the distance estimation task because more than half of their trials fell outside the ±2.5 SD range (participants P48 and P49). These participants are retained in `lane_change_dataset.csv` and `participant_metadata.csv`, as the exclusion criterion was specific to the distance estimation task. Two additional participants contained isolated outlying trials (one and two trials, respectively); these isolated outlying trials were replaced using mean-value imputation within the corresponding dependent variable. Following preprocessing, the distance estimation dataset contains 56 participants and the lane-change dataset contains 58 participants. --- ## Analysis scripts R scripts for the analyses reported in the associated manuscript. All scripts expect the dataset CSV files in the same folder as the scripts, or edit `data_dir` at the top of each script. Requirements: R >= 4.0. Each script checks for and installs its required packages (lme4, lmerTest, emmeans, tidyverse, performance, ggplot2, ggsignif, dplyr, rmcorr) from CRAN if missing. Run order: 1. **CMS_continuous_age_analysis.R** - Primary analyses. Linear mixed-effects models with driver age as a continuous predictor for signed distance estimation error (DistErr), relative error (RelErr), and time-to-contact (TTC), including simple-slope (estimated marginal trend) analyses, the quadratic age check, the gap-split sensitivity analysis, and model-based predictions with confidence intervals. Writes three CSV files of model predictions to the working directory. 2. **Within_subject_figures.R** - Publication figures for the within-subject effects. Must be run in the same R session after script 1, as it uses the emmeans objects created there. Figure export lines (`ggsave`) are provided but commented out. 3. **confidence_analysis.R** - Confidence rating analyses: condition effects via linear mixed-effects models, repeated-measures correlations (rmcorr) between trial-level confidence and objective performance, and participant-level Spearman correlations. Self-contained; can be run independently. 4. **unsigned_error_analysis.R** - Unsigned (absolute) error analyses for the distance estimation task, using the same model structure as the primary analyses, with partial eta squared computed from the F-based approximation used in the manuscript. Self-contained; can be run independently. --- ## Ethics and anonymisation The shared datasets do not contain directly identifying personal information. Participant IDs are anonymised. Raw interview recordings, interview transcripts, and any potentially identifying qualitative material are not included in the shared repository for ethical and privacy reasons. --- ## Suggested citation Pifferi, G., Thulinsson, F., Söderlund, N., Brunnström, K., Rafiei, S., Schenkman, B., Djupsjöbacka, A., Sperandio, I., & Andrén, B. (2026). Dataset for: *Effects of Camera Height, Field of View, and Driver Age on Depth Judgement and Lane-Change Decisions in Camera Monitor Systems* [Data set]. Zenodo. DOI: https://doi.org/10.5281/zenodo.20055125