Tsui, Claire · Briga, Michael · Komdeur, Jan · et al.
hybrid · semantic + lexical · 2438 datasets ranked · 1.77s
Tsui, Claire · Briga, Michael · Komdeur, Jan · et al.
Data and code for analysis in manuscript titled "Asynchrony of ageing among traits in a wild bird population" Dataframes ending with"_28_5.csv" and "survival_model.csv" are used in scripts model1-13, of which the output is plotted using "new model outputs.R" Code for Figures 1 and 2 are in script "new model outputs.R" asymmetry bivar ver3.R runs the bivariate models used to estimate the degree of synchrony of ageing. scripts starting with "aic.." are used in the analysis for age by lifespan interaction
Bell, Peter
2 files · 14 MB · pdf, zipdeclared
This working paper asks how Canada should assess strategic resource infrastructure that may not earn a commercial return on its own. A road, port, power system, pipeline, smelter, or processing plant can appear uneconomic when evaluated as an individual asset while still enabling new production, preserving difficult-to-replace processing capacity, supporting several users, or improving security of supply. The paper develops a framework for deciding when public support for such an asset may be justified. Its central rule is that an asset-level loss is defensible only when it produces wider benefits that are specific, measurable, and subject to effective public oversight. The analysis draws on Canadian wartime industrial mobilization, concentration in critical-mineral supply chains, proposed support for processing capacity at Trail, current infrastructure and northern development programs, and cautionary cases involving mining subsidies, managed decline, remote transport, and major-project governance. The paper converts the argument into a twelve-question approval test and a measurement framework for tracking public cost, avoided closure, new production, secure supply, shared infrastructure use, and signs of failure. It does not recommend a particular project and does not argue that all loss-making assets deserve public support. Its purpose is to distinguish infrastructure that creates durable public value from subsidy, bailout, or white-elephant risk.
Chen, Hang
1 files · 151 MB · zipdeclared
What's Changed Add GPT API-based multi-agent system for automated ERT workflows by @geohang with @Copilot in https://github.com/geohang/PyHydroGeophysX/pull/3 Add multi-LLM provider support and position agent system for cross-modal geophysics by @geohang with @Copilot in https://github.com/geohang/PyHydroGeophysX/pull/4 Add ClimateDataAgent for PyDaymet integration with ERT workflows by @geohang with @Copilot in https://github.com/geohang/PyHydroGeophysX/pull/5 Add ClimateDataAgent integration for cross-modal climate-ERT reasoning by @geohang with @Copilot in https://github.com/geohang/PyHydroGeophysX/pull/6 New Contributors @geohang with @Copilot made their first contribution in https://github.com/geohang/PyHydroGeophysX/pull/3 Full Changelog : https://github.com/geohang/PyHydroGeophysX/compare/1.0...v0.3.0
Tabiri, Josephine Konamah
1 files · 11 MB · zipdeclared
This repository contains photographic images from six stores used in my thesis reserach. The files document the visual material analyzed in the study.
Levine, Jonathan · MacEachern, Leonard
1 files · 510 KB · zipdeclared
Supplementary artifact for the IEEE Embedded Systems Letters manuscript "Threshold-Based Batch Normalization for Integer-Only Binarized Neural Network Inference" by Jonathan Levine and Leonard MacEachern. The package contains FPGA HLS sources, exported trained-network parameters, ECDF/statistics data, TCL reproduction scripts, selected logs and reports, and helper scripts for the MNIST, CERN jet-substructure tagging, and UNSW-NB15 workloads. It supports the paper's threshold-export formulation for integer-only dense BNN inference and the Alveo U250 HLS results reported in the manuscript. Repository: https://github.com/maceacla/tbbn-bnn-hls
Anonymus
1 files · 424 KB · zipdeclared
Analysis code for the manuscript "Bayesian sensory integration explains ball-count bias in Major League Baseball umpires" (submitted to Communications Psychology; author information withheld for double-anonymised peer review). This archive contains the complete analysis pipeline for quantifying the ball-count-dependent strike/ball decision bias of MLB home-plate umpires and explaining it with a Bayesian sensory integration model, together with the derived result files and figures reported in the manuscript. Data: MLB 2015-2024 regular and post-season games, 451,172 called pitches (four-seam fastballs thrown by right-handed pitchers to right-handed batters). Pipeline (numbered folders 00-05): - Data acquisition (Statcast pitch-tracking data and home-plate umpire assignments) - Merging and inclusion-criteria filtering - Count-wise psychometric (probit) function fits (point of subjective equality and perceptual uncertainty) - Two-component Gaussian Mixture Model of called-pitch locations (count-specific priors) - Trial-level Bayesian sensory integration model fit (count-specific vs. universal prior) - Umpire-wise individual-differences analysis (n = 89) Environment: Python 3.13 (pybaseball 2.2.7, pandas, pyarrow) and MATLAB R2022b (Statistics and Machine Learning Toolbox). Raw and intermediate pitch-tracking data are not included, owing to file size and to avoid redistributing Baseball Savant data; they can be regenerated with the scripts in steps 00-01. See the included README.md for the full folder structure, reproduction instructions, and the correspondence between output files and the values reported in the manuscript.
Bradley, Alex
1 files · 244 KB · zipdeclared
Companion code for fitting hierarchical Bayesian spatial models that calibrate sedimentary leaf-wax n-C29 alkane hydrogen isotope ratios against precipitation isotope composition and environmental covariates. Implements 14 model variants in Stan with a Matern 3/2 Gaussian process predictive-process approximation on 125 globally distributed knots. Compares spatial and non-spatial calibrations across n = 1,128 surface sediment and soil sites from 73 publications.
Madhusudan, Gujral
6 files · 29 MB · parquetdeclared
Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.
Krajnc, Matej · Comi, Troy · Miao, Siqi · et al.
10 files · 25 GB · csv, zipdeclared
This dataset accompanies the manuscript "A Controlled in Silico Benchmark for GNN Prediction of Tissue Dynamics." It contains model prediction outputs, trained checkpoints, train/validation/test splits, spring-embedding outputs, generated analysis figures, analysis tables, and manuscript-specific diagnostic outputs used to reproduce the post-prediction analyses and figures. The dataset is distributed as logical ZIP archives with file-level and archive-level SHA-256 checksums. For questions, contact Tomer Stern at tomers@umich.edu.
Garrett, Patrick T. · Yates III, John R.
1 files · 83 MB · zipdeclared
Added Precursor-space MS1 gates ( tdfpy.noise.gates ). Two acquisition-aware NoiseFilter s that drop MS1 signal the instrument never fragments — signal that cannot become an identification. Both convert their (m/z, 1/K0) region once to per-scan integer TOF-index intervals (via the run calibration) and test membership with a vectorised binary search; both no-op (keep everything) when the run carries no region. Ported from the dnoise Rust tool. Compose like any filter, e.g. noise=[SelectionPolygonGate(), MadThreshold(k=3)] . SelectionPolygonGate (ddaPASEF) — keeps only MS1 points inside the run's PASEF selection polygon (the "IMS PolygonFilter" read from analysis.tdf 's GroupProperties ). A generalisation of ChargeStateRegion from a single line to the real acquisition polygon. Skipped on diaPASEF (where the same property stores window quads). Padded in physical units ( mz_pad default 5 Da, im_pad default 0.05 1/K0) so an edge precursor keeps its isotopic envelope / mobility spread rather than being clipped at a hard polygon boundary; pass mz_pad=0.0, im_pad=0.0 for a hard cutoff. DiaMs1WindowGate (diaPASEF) — keeps only MS1 points inside the union of the isolation windows ( DiaFrameMsMsWindows ); everything outside is a precursor the method never isolates. Windows are padded in physical units ( mz_pad default 5 Da, im_pad default 0.05 1/K0). No-op on ddaPASEF. Both gates now no-op (keep everything) on non-MS1 frames rather than testing fragment peaks against the MS1 precursor region (which would empty an MS2 spectrum), so they are safe to leave in a noise=[...] list applied across frame types. build_window_intervals clamps a negative scan_lo to 0 and skips boxes lying wholly outside the scan range instead of letting a negative bound wrap around via Python indexing.
Fabbri, Renato
1 files · 3.5 MB · zipdeclared
BSC Lab is an open-source sensory stimulation platform with two integrated layers: a precision multi-engine audiovisual stimulation application, and a knowledge layer built on the SSTIM ontology (OWL class hierarchy, multilingual SKOS vocabulary, SHACL validation shapes, exposure module, and external alignments) published at https://w3id.org/sstim under CC BY 4.0. Each tagged release is archived here for citation.
MENEGUZZO, FRANCESCO · Zabini, Federica
4 files · 146 KB · csvdeclared
This record contains the de-identified, analysis-ready dataset and Python analysis code associated with a school-based quasi-experimental pilot study on therapist-guided forest therapy and persistent anxiety symptoms in adolescents. The study involved two fourth-year high-school classrooms in Cecina, Tuscany, Italy: one intervention classroom that attended four therapist-guided forest therapy sessions in a coastal pine forest, and one control classroom that followed usual school activities. The dataset includes anonymized student codes, classroom allocation, SCAS total raw scores and derived reduction scores across repeated assessments, POMS-A acute mood-state variables for the intervention classroom, exploratory post-intervention nature-exposure and connectedness variables, and exploratory school-performance variables. The repository includes four files: 1. forest_therapy_adolescent_anxiety_dataset.csv: de-identified analysis-ready dataset. 2. README_forest_therapy_adolescent_anxiety_dataset.txt: dataset metadata and variable descriptions. 3. analyze_forest_therapy_adolescent_anxiety.py: Python script used to reproduce the main analyses and generate analysis outputs. 4. README_analyze_forest_therapy_adolescent_anxiety.txt: operational guide for running the analysis script and interpreting the output files. The dataset uses semicolon-separated values. Missing values are encoded as NaN and must not be interpreted as zero. Because the study involved minors, item-level questionnaire responses and more granular school records are not shared; only de-identified and analysis-ready variables compatible with privacy, ethical approval, and consent constraints are provided.
Bradley, Alex
1 files · 12 MB · zipdeclared
R package for reconstructing precipitation hydrogen isotope ratios from sedimentary leaf-wax n-C29 alkane measurements using the spatially-aware hierarchical Bayesian calibration of Bradley (2026). Inverts the forward calibration with full uncertainty propagation, combining analytical error, residual variance, slope posterior, and the spatial Gaussian process intercept. Supports per-record change detection with autocorrelation-adjusted thresholds and a four-level claim taxonomy from measurement-level changes through uniquely- attributable precipitation-isotope claims.
Azziz, Ricardo
1 files · 1.6 MB · docxdeclared
Supplemental Table 1. Meeting Agenda of the 2024 PCOS Challenge-CDC Stakeholder Meeting on Testosterone Reference Interval Standardization. August 26, 2024. Centers for Disease Control and Prevention, Atlanta, Georgia. Supplemental Table 2. Proposed Four-Stage Workflow for Developing Standardized Testosterone Reference Intervals Using Existing Study Data.
Juniper L. Simonis
1 files · 86 MB · zipdeclared
Tools for interacting with the publicly available California Delta Fish Salvage Database, including continuous deployment of data access, analysis, and presentation.
Austin, Timothy · do Valle Chagas Azaneu, Marina · Roughan, Moninya
26 files · 29 MB · netcdf, pdf, pngdeclared
Data collected from a temperature mooring at Lord Howe Island maintained by UNSW Sydney and funded by Parks Australia. The mooring position is longitude = 158.97°E and latitude = -31.51°, and local depth of approximately 52 m. The data were sampled using a series of thermistors (aqualogger 520PTs) deployed on a mooring line at 4m intervals through the water column, with shallowest instrument at 13 m and deepest at 53 m. The time period spans between 14-05-2025 and 22-04-2026. IMOS standard data quality assurance and quality control processes have been followed and the data formatted following IMOS conventions. Data quality control includes automated routines and visual inspection (expert QC) and flagging of obvious errors. File are c.f. compliant NetCDF files, and file name format follows IMOS conventions and includes sampling period in the format: UNSW_Lord_Howe_Marine_Park_TZ_ yyyymmddThhmmss Z_LH050_FV01_ LH050-2511-Aqualogger-AQUAlogger-520PT16-max160m-13_END- yyyymmddThhmmssZ.
Padilla-Villanueva, Johel
1 files · 457 KB · zipdeclared
Academy Learning Tau is an open educational and research software platform that operationalizes the Systemic Tau framework and the Discrete Extramental Clock (RECD) for complex time-series analysis. It provides ordinal metrics (τ_s, nested RECD levels), classical early-warning signals for dual reading, surrogate null models, multilingual pedagogy (Spanish, English, French), and a reproducible Streamlit laboratory for teaching and exploratory research.
Aouladali, Amal · Alaoui, Souad · Hnini, Abdelhalim
4 files · 491 KB · zipdeclared
Replication code and measured artifacts for KeystrokeAuthChain: a keystroke-dynamics continuous-authentication Transformer evaluated on the public CMU Killourhy-Maxion dataset (EER 7.47%), coupled to an on-chain audit layer that anchors each authentication decision on the Ethereum Sepolia testnet for non-repudiation and third-party-verifiable auditing. Includes the recognition pipeline, paired baselines, significance tests, the audit smart contracts and gas measurements, and the real end-to-end anchoring run (10 genuine decisions, on-chain decisionHash verified 10/10). The CMU dataset is not redistributed here (see data/README.md inside the archive).
Padilla-Villanueva, Johel
1 files · 356 KB · zipdeclared
Plataforma web educativa y de investigación para el paradigma Tau Sistémica y el Reloj Extramental Discreto (RECD). Incluye un laboratorio interactivo para el análisis de señales biológicas, ecológicas y financieras mediante ordinal patterns, EWS y TDA.
Bradley, Alex
1 files · 10 MB · zipdeclared
Posterior draws from 14 hierarchical Bayesian leaf-wax-to-precipitation calibration models, fit in Stan on the frozen calibration run c2_run_20260626 (n = 1128 observations; Bradley 2026, Communications Earth and Environment). Each file is a posterior::draws_df with 1000 stratified draws across 8 MCMC chains. Spatial models include 125 per-knot latent z values on a Fibonacci sphere lattice. Used by the leafwax R package as the full-resolution backing data; the package ships a 100-draw fixture and downloads the full posteriors from this archive on first use. Supersedes the v1.x deposit (v10 / n = 1129 lineage).