Comalada i Pla, Francesc
hybrid · semantic + lexical · 290 datasets ranked · 9.62s
1 files · 224 KB · docx
Dataset containing the systematic review matrix and extracted variables supporting the article "Bridging ecological restoration and social legitimacy: a systematic review of Cultural Ecosystem Services in inland aquatic ecosystems", accepted for publication in People and Nature.
Lopez, Annalaura · Greco, Margherita · Marcolli, Beatrice · et al.
4 files · 17 KB · docx, xlsx
This dataset originates from a study aiming to valorise Ciuta sheep, a local breed native from the Italian Central Alps, through the characterization of nutritional quality and chemical composition of fresh meat (loins) and one traditional dry-cured product. Specifically, the research focused on determining the chemical composition of Ciuta sheep meat and on identifying key changes in its chemical profile during dry curing process, hypothesizing that such chemical fingerprint may suggest some markers linked to the production system, geographical origin, and traditional processing techniques. For this reason, for bthe dry-cured product, both an aliquot of fresh meat before and after transformation and dry-curing was sampled and analysed. Regarding loins, three commercial categories (lambs, hoggets and mutton) were considered, in order to define any possible difference induced by age of the sheep (and physiological factors, such as rumen development). The dataset includes chemical data regarding the proximate composition (moisture, protein, fat, ash, salt content for the dry-cured product) and energy content of fresh and dry-cured meat; the fatty acids content of fresh and dry-cured meat product; the volatile profile of fresh and dry-cured meat product. Results from analysis performed in our study suggested that the development of high-quality dry-cured products could provide a strategy to valorise Ciuta sheep meat, especially from adult animals (culled ewes and rams), while fresh meat production could focus on lambs. The complex volatile profile detected was influenced by both the farming system and traditional processing methods.
Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.
2,095 rows · 907 KB · fasta, tsv
A virus genome database representing 21,015 near-complete virus genomes collected from untargeted ultra-deep RNA/DNA combined sequencing of wastewater. Sequence data was provided by the CASPER consortium and raw data may be found on NCBI SRA under bioprojects PRJNA1247874 and PRJNA1198001. Data underwent read trimming, rRNA and human read removal, de novo assembly, and selection of high-quality viral contigs. Contigs were clustered at 95% identity and 85% query coverage to dereplicate. Chimera-checking required at least two independent assemblies of the same viral genome or presence of the genome in another reference database. Annotation made use of RdRpCATCH, geNomad, checkV, BLASTN against NCBI core-nt, and RNAVirHost. The RdRp fasta files contain representative RdRp sequences identified through homology to major RdRp reference databases and clustered at 90% sequence identity over 75% sequence coverage. Included sequences contain all three conserved RdRp motifs (A, B, and C) arranged in either the canonical ABC configuration or the permuted CAB configuration.
Hubbard, Alfred · Solares, Edwin · Hemming-Schroeder, Elizabeth
88 rows · 18 KB · fasta
These are the files needed to run the Broad Institute's malaria amplicon pipeline for the PvGAP Plasmodium vivax panel, described in detail here . They consist of FASTA files containing the forward and reverse primers and another FASTA file containing reference sequences for each target, derived from the PvP01 reference genome.
Yuan, Guangyuan
72 rows × 13 cols · 6.0 KB · csv, fasta
11 numeric · 2 categorical
This dataset supports the findings of the manuscript "Root anatomical traits modulate the assembly and nitrogen transformation potential of root-associated microbiomes in a temperate steppe" (NPH-MS-2026-55667). It contains root traits data, bacterial 16S rRNA gene absolute abundances, functional genes relative abundances, DNA extraction metadata, and phylogenetic marker sequences for 37 plant species from a temperate steppe ecosystem. The dataset includes the following files: 1. root traits.csv - Root traits including average diameter (AD), specific root length (SRL), specific root area (SRA), root tissue density (RTD), root nitrogen content (RNC), root carbon content (RCC), carbon‑nitrogen ratio (RCN), cortex layer number (CLN), cortex thickness (CT), and the ratio of cortex thickness to root diameter (CTRD). The first column lists plant species names. 2. Absolute abundance of 16S rRNA gene.csv - Quantitative PCR (qPCR) derived absolute abundances of bacterial 16S rRNA gene copies (copies/ng DNA) across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 3. DNA extraction sample weight.csv - Fresh weight (grams) of root material used for DNA extraction for each sample, linked by SampleID to the abundance data. 4. DNA extraction concentration.csv - Qubit‑measured DNA concentrations (ng/μL) and the sample volume (μL) used for quality control, together with sample metadata. 5. 37species.fasta - DNA sequences of two chloroplast markers (matK and rbcL) for the 37 plant species included in the study. The sequences are in FASTA format with headers formatted as ">Species". These were used for host phylogeny construction and Pagel's λ analyses. 6. Quantitative PCR results of functional gene.csv - Quantitative PCR (qPCR) derived relative abundances of bacterial 16S rRNA gene and functional genes across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 7. README.md - A detailed description of each file, column headers, abbreviations, units, and any missing value codings (NA). All data are provided to ensure transparency and reproducibility of the analyses. For methodological details, please refer to the Materials and Methods section of the associated publication. These data are under embargo until the associated research article is published. After that date, they will be freely available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license. During the embargo period, the metadata (title, authors, abstract) and the DOI remain publicly visible, but the data files are not accessible. For access requests before the embargo expires, please contact the corresponding author.
Budzinski, Lisa · Beenken, Anne Elisabeth · Sempert, Toni · et al.
9 rows × 1 cols · 743 B · csv, docx, zip
1 categorical
We have investigated an IgG4-RD (IgG4-RD) cohort by our multi-parameter microbiota flow cytometry approach to characterise the microbiota on single-cell level for attributes of the disease. The microbiota is isolated from stool samples and stained according to the published protocol for (a) host immunoglobulins IgA1, IgA2, IgM, IgG and (b) agglutinin binding to mannose, galactose or N-Acetyl-glucosamine surface sugar moieties. For all samples we also determined the microbiome composition by 16S rRNA (V3-V4) sequencing on the illumina MiSeq platform. We provide the raw .fcs and FASTQ files of 40 IgG4-RD patients. For comparison we additionally analysed 36 healthy donors. All .fcs files were generated on BD Influx®. The metadata is collected in the provided meta.csv. The staining parameters are summarized in provided panel.csv.
Li, Zhiyao · Wang, Ningbo · Zhong, Jiahao
6 files · 48 MB · docx, zip
GIFT-BDS is a regional ionospheric total electron content (TEC) and TEC-gradient dataset over China derived from BeiDou geostationary Earth orbit (GEO) observations and a dense ground-based GNSS receiver network. The dataset is designed to provide high-resolution observations of ionospheric TEC variability and horizontal TEC-gradient structures over China and adjacent regions. The versioned release covers the period from 19 July 2024 to 31 December 2025, corresponding to DOY 201 of 2024 to DOY 365 of 2025. The geographical coverage is 15°N-50°N and 95°E-135°E. The dataset is provided in daily NetCDF files and contains two product levels. Level-1 products provide observation-level GEO-derived slant TEC (STEC) and rate of TEC index (ROTI) records for individual receiver-GEO satellite lines of sight, with a temporal resolution of 30 s. Level-2 products provide gridded regional TEC and TEC-gradient variables, including VTEC, VTEC t , ROTI, GIX, GIX std , GIX x , GIX y , GIX t,x , and GIX t,y , with a temporal resolution of 15 min. IPP-based variables are provided on a 1° × 1° grid, while inter-IPP-gradient variables are provided on a 0.25° × 0.25° grid. The main processing steps include observation screening, cycle-slip and data-gap detection, continuous-arc segmentation, carrier-to-code leveling, satellite and receiver DCB correction, IPP calculation, inter-IPP pair selection, gradient estimation, and gridding. Quality control is applied before release. Missing values may occur because of station outages, data gaps, quality-control exclusions, or insufficient valid samples within a grid cell. Users should check the NetCDF variable attributes, including units and fill values, before analysis. The dataset is suitable for regional ionospheric studies, TEC-gradient monitoring, space-weather-related analyses, and investigations of ionospheric effects on GNSS positioning applications.
Jungmyoung, Son · Sihoon, Lee · Jiyeon, Hong
3 files · 20 KB · docx, xlsx
This dataset provides the complete double-coding matrix, PRISMA 2020 checklist, and search strategy supporting the systematic review "Representation-to-AI Transformation in K-12 Generative AI Learning: A Theory-Building Systematic Review of Semantic Transformation Mechanisms." It includes: (1) study-level tier classification (Core/Supporting/Context) for two independent coders and consensus tier for all 18 included studies; (2) the full semantic transformation unit (STU) coding matrix (18 studies x 10 STUs = 180 cells) with pre-consensus and consensus scores; (3)evidence-weighting consensus scores; (4) inter-rater reliability statistics (Cohen's kappa); (5) the completed PRISMA 2020 checklist; and (6) the full database-specific Boolean search strategy.
Requena Rolanía, Jose María · Greif, Gonzalo · ROBELLO, CARLOS
1 files · 8.0 MB · fasta
This dataset contains the genome sequence for Trypanosoma cruzi (strain Dm28c). This genome sequence was de novo assembled using PacBio Hi-Fi and Illumina sequencing platforms by Greif et al (2026. PMID: 41501640). The genome was assembled into 32 contigs, which represent complete chromosomes. The provided Fasta file also contains an additional contig corresponding to the maxicircle (mitochondrial genome) sequence. The Fasta files included in this dataset were downloaded from GenBank (assembly GCA_044048535.1; May 22, 2026). Additional information about the Dm28cT2T genome assembly and gene annotations may be accessed through the link: https://cruzi.pasteur.uy/
Burman, Nathaniel · Buyukyoruk, Murat · Wiegand, Tanner · et al.
4 files · 8.0 MB · fasta
This folder contains a multiple sequence alignment of Cas7 homologs in .fasta format, the domain-level annotations from PFAM and CasFinder, and an associated phylogenetic tree in .newick format.
Yang, Qingliu
6 files · 8.0 MB · bzip2
Dataset Description This dataset contains 3D lightning location results, DALMA and FALMA waveform for two Energetic Compact Stroke (ECS) events. location results are included: HF3D_1732785151.dat - 3D lightning locations for the ECS leader A flash. HF3D_1734785454.dat - 3D lightning locations for the ECS leader B flash. The timestamp 1734785454 and 1732785151 corresponds to the occurrence time of the lightning flash in Japan Standard Time. File format and parameters The first row contains the lightning occurrence time. Column descriptions: Time (ms) - time relative to the lightning source. X, Y, Z (m) - 3D spatial coordinates relative to ground level. The origin (0, 0, 0) corresponds to latitude 36.76°N and longitude 136.76°E. FALMA and DALMA waveform ECSLeaderA_DALMA_waveform.bz2 is DALMA waveform of Leader A. ECSLeaderA_FALMA_waveform.bz2 is FALMA waveform of Leader A. ECSLeaderB_DALMA_waveform.bz2 is DALMA waveform of Leader B. ECSLeaderB_FALMA_waveform.bz2 is FALMA waveform of Leader B. This dataset allows analysis of the spatial and temporal development of these two ECS flashes.
Kovaliov, Michael
1 files · 100 MB · parquet
Madhusudan, Gujral
6 files · 29 MB · parquetdeclared
Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.
Azziz, Ricardo
1 files · 1.6 MB · docxdeclared
Supplemental Table 1. Meeting Agenda of the 2024 PCOS Challenge-CDC Stakeholder Meeting on Testosterone Reference Interval Standardization. August 26, 2024. Centers for Disease Control and Prevention, Atlanta, Georgia. Supplemental Table 2. Proposed Four-Stage Workflow for Developing Standardized Testosterone Reference Intervals Using Existing Study Data.
García Mina, Janeth Elizabeth · Rodríguez Garófalo, Napoleón Hernán · Rumbaut-Rangel, Dayron · et al.
6 files · 737 KB · docx, pdfdeclared
El objetivo de este estudio es evaluar la importancia de la inteligencia artificial (IA) en el proceso de formación académica de estudiantes de educación básica, específicamente en la enseñanza de ciencias naturales. La investigación destaca cómo la IA puede personalizar el aprendizaje, adaptándose a los estilos individuales de cada alumno, lo que resulta en una experiencia educativa más efectiva y significativa. Se utilizó un diseño cuasiexperimental con una muestra de 60 estudiantes de sexto año, distribuidos en un grupo control (30 estudiantes) que recibió instrucción tradicional y un grupo experimental (30 estudiantes) que utilizó herramientas de IA, como chatbots educativos y plataformas de aprendizaje adaptativo. Los resultados mostraron que el 70% de los estudiantes del grupo experimental alcanzaron calificaciones superiores a 8, en comparación con solo el 30% del grupo control. Este hallazgo sugiere que la implementación de herramientas de IA puede transformar la educación básica, mejorando tanto el rendimiento académico como la motivación de los estudiantes. En conclusión, el estudio resalta la necesidad de seguir investigando y desarrollando nuevas estrategias educativas que integren la IA para maximizar su potencial en la educación.
Makhado, Langanani Christinah · Raliphaswa, Ndidzulafhi Selina
9 files · 157 KB · docxdeclared
This dataset contains anonymized transcripts derived from face-to-face interviews conducted with pregnant women and breastfeeding mothers as part of the research study. The interviews were audio-recorded and transcribed verbatim to capture participants' experiences, perceptions, and insights related to the study topic. All personally identifiable information has been removed to protect participant confidentiality. The transcripts are provided to support transparency, reproducibility, and further research related to the findings reported in the associated article
Fontelo, Paul
1 files · 26 KB · docxdeclared
Auto-Brewery Syndrome (ABS), also known as gut fermentation syndrome, is a condition in which microbial fermentation of dietary carbohydrates generates endogenous ethanol, causing signs and symptoms of alcohol intoxication without alcohol consumption. Misdiagnosis as alcohol use disorder (AUD) is common and carries severe medical, psychosocial, legal, and forensic consequences. The objective of thus narrative review is to synthesize the clinical literature on ABS for practicing clinicians and to present the first systematic analysis of ABS susceptibility in Asian populations - including communities of Asian descent in Western countries.
Hackl, Jürgen
6 files · 132 MB · parquet, tiffdeclared
EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).
Intergovernmental Science-Policy Platform on Biodiversity and Ecosystem Services
2 files · 1.5 MB · docx, pdfdeclared
This is the report on the indigenous and local knowledge (ILK) dialogue workshop for the first order draft of the summary for policymakers and the second order draft of the IPBES assessment of the sustainable use of wild species (the "sustainable use assessment"). It was held from 17-21 May 2021, online, due to the ongoing COVID-19 pandemic. The report aims to provide a written record of the dialogue workshop, which can be used by assessment authors to inform their work on the sustainable use assessment, and also by all dialogue participants who may wish to review and contribute to the work of the assessment moving forward. The report is not intended to be comprehensive or give final resolution to the many interesting discussions and debates that took place during the workshop. Instead, it is intended as a written record of the discussions, and this conversation will continue to evolve in the course of the assessment process. For this reason, clear points of agreement are discussed, but diverging views among participants are also presented for further attention and discussion.
Ogah, Odey
4 files · 3.4 MB · docx, pdf, zipdeclared
The study investigated the determinants of the migration of youth in Benue State from agricultural activities using a multi stage survey research design. The population of the study comprised 90 households in the three zones in Benue state selected by simple random sampling. A sampling proportion of 5% was used and a sample size of 270 was drawn. The instrument for data collection was a self-developed questionnaire. Data collected was analysed using frequencies distribution, mean, standard deviation and probit regression to achieve the research objectives while chi square goodness of fit was used to test the hypothesis. The findings revealed that 69.0% of the youth involved in agricultural activities in the study area were within the ages of 26- 30 years old while few (7.7%) are of 31years and above. The findings of the study revealed that land tenure practices have high influence on agricultural land developmental activities of youth in the study area. The study found that there is a plethora of factors responsible for seasonal migration of youth in the study area and these includes; Lack of access to modern farming equipment and technology; limited economic opportunities in the agricultural sector. The study concludes that youth migration affects agricultural activities in the study area. The study therefore recommended among others, skills diversification programme that will equip youth in diverse skills beyond seasonal agricultural works that could be utilised all year round. Year round employment opportunities that would provide consistent employment for the youth, education and training initiatives that would focus on agricultural techniques and sustainable practices, access to financial resources and collaboration and networking between local governments, NGOS, and the private sectors.