Bell, Peter
2 files · 14 MB · pdf, zipdeclared
hybrid · semantic + lexical · 1959 datasets ranked · 2.56s
Bell, Peter
2 files · 14 MB · pdf, zipdeclared
This working paper asks how Canada should assess strategic resource infrastructure that may not earn a commercial return on its own. A road, port, power system, pipeline, smelter, or processing plant can appear uneconomic when evaluated as an individual asset while still enabling new production, preserving difficult-to-replace processing capacity, supporting several users, or improving security of supply. The paper develops a framework for deciding when public support for such an asset may be justified. Its central rule is that an asset-level loss is defensible only when it produces wider benefits that are specific, measurable, and subject to effective public oversight. The analysis draws on Canadian wartime industrial mobilization, concentration in critical-mineral supply chains, proposed support for processing capacity at Trail, current infrastructure and northern development programs, and cautionary cases involving mining subsidies, managed decline, remote transport, and major-project governance. The paper converts the argument into a twelve-question approval test and a measurement framework for tracking public cost, avoided closure, new production, secure supply, shared infrastructure use, and signs of failure. It does not recommend a particular project and does not argue that all loss-making assets deserve public support. Its purpose is to distinguish infrastructure that creates durable public value from subsidy, bailout, or white-elephant risk.
Chen, Hang
1 files · 151 MB · zipdeclared
What's Changed Add GPT API-based multi-agent system for automated ERT workflows by @geohang with @Copilot in https://github.com/geohang/PyHydroGeophysX/pull/3 Add multi-LLM provider support and position agent system for cross-modal geophysics by @geohang with @Copilot in https://github.com/geohang/PyHydroGeophysX/pull/4 Add ClimateDataAgent for PyDaymet integration with ERT workflows by @geohang with @Copilot in https://github.com/geohang/PyHydroGeophysX/pull/5 Add ClimateDataAgent integration for cross-modal climate-ERT reasoning by @geohang with @Copilot in https://github.com/geohang/PyHydroGeophysX/pull/6 New Contributors @geohang with @Copilot made their first contribution in https://github.com/geohang/PyHydroGeophysX/pull/3 Full Changelog : https://github.com/geohang/PyHydroGeophysX/compare/1.0...v0.3.0
Tabiri, Josephine Konamah
1 files · 11 MB · zipdeclared
This repository contains photographic images from six stores used in my thesis reserach. The files document the visual material analyzed in the study.
Levine, Jonathan · MacEachern, Leonard
1 files · 510 KB · zipdeclared
Supplementary artifact for the IEEE Embedded Systems Letters manuscript "Threshold-Based Batch Normalization for Integer-Only Binarized Neural Network Inference" by Jonathan Levine and Leonard MacEachern. The package contains FPGA HLS sources, exported trained-network parameters, ECDF/statistics data, TCL reproduction scripts, selected logs and reports, and helper scripts for the MNIST, CERN jet-substructure tagging, and UNSW-NB15 workloads. It supports the paper's threshold-export formulation for integer-only dense BNN inference and the Alveo U250 HLS results reported in the manuscript. Repository: https://github.com/maceacla/tbbn-bnn-hls
Anonymus
1 files · 424 KB · zipdeclared
Analysis code for the manuscript "Bayesian sensory integration explains ball-count bias in Major League Baseball umpires" (submitted to Communications Psychology; author information withheld for double-anonymised peer review). This archive contains the complete analysis pipeline for quantifying the ball-count-dependent strike/ball decision bias of MLB home-plate umpires and explaining it with a Bayesian sensory integration model, together with the derived result files and figures reported in the manuscript. Data: MLB 2015-2024 regular and post-season games, 451,172 called pitches (four-seam fastballs thrown by right-handed pitchers to right-handed batters). Pipeline (numbered folders 00-05): - Data acquisition (Statcast pitch-tracking data and home-plate umpire assignments) - Merging and inclusion-criteria filtering - Count-wise psychometric (probit) function fits (point of subjective equality and perceptual uncertainty) - Two-component Gaussian Mixture Model of called-pitch locations (count-specific priors) - Trial-level Bayesian sensory integration model fit (count-specific vs. universal prior) - Umpire-wise individual-differences analysis (n = 89) Environment: Python 3.13 (pybaseball 2.2.7, pandas, pyarrow) and MATLAB R2022b (Statistics and Machine Learning Toolbox). Raw and intermediate pitch-tracking data are not included, owing to file size and to avoid redistributing Baseball Savant data; they can be regenerated with the scripts in steps 00-01. See the included README.md for the full folder structure, reproduction instructions, and the correspondence between output files and the values reported in the manuscript.
Bradley, Alex
1 files · 244 KB · zipdeclared
Companion code for fitting hierarchical Bayesian spatial models that calibrate sedimentary leaf-wax n-C29 alkane hydrogen isotope ratios against precipitation isotope composition and environmental covariates. Implements 14 model variants in Stan with a Matern 3/2 Gaussian process predictive-process approximation on 125 globally distributed knots. Compares spatial and non-spatial calibrations across n = 1,128 surface sediment and soil sites from 73 publications.
Madhusudan, Gujral
6 files · 29 MB · parquetdeclared
Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.
Krajnc, Matej · Comi, Troy · Miao, Siqi · et al.
10 files · 25 GB · csv, zipdeclared
This dataset accompanies the manuscript "A Controlled in Silico Benchmark for GNN Prediction of Tissue Dynamics." It contains model prediction outputs, trained checkpoints, train/validation/test splits, spring-embedding outputs, generated analysis figures, analysis tables, and manuscript-specific diagnostic outputs used to reproduce the post-prediction analyses and figures. The dataset is distributed as logical ZIP archives with file-level and archive-level SHA-256 checksums. For questions, contact Tomer Stern at tomers@umich.edu.
Garrett, Patrick T. · Yates III, John R.
1 files · 83 MB · zipdeclared
Added Precursor-space MS1 gates ( tdfpy.noise.gates ). Two acquisition-aware NoiseFilter s that drop MS1 signal the instrument never fragments — signal that cannot become an identification. Both convert their (m/z, 1/K0) region once to per-scan integer TOF-index intervals (via the run calibration) and test membership with a vectorised binary search; both no-op (keep everything) when the run carries no region. Ported from the dnoise Rust tool. Compose like any filter, e.g. noise=[SelectionPolygonGate(), MadThreshold(k=3)] . SelectionPolygonGate (ddaPASEF) — keeps only MS1 points inside the run's PASEF selection polygon (the "IMS PolygonFilter" read from analysis.tdf 's GroupProperties ). A generalisation of ChargeStateRegion from a single line to the real acquisition polygon. Skipped on diaPASEF (where the same property stores window quads). Padded in physical units ( mz_pad default 5 Da, im_pad default 0.05 1/K0) so an edge precursor keeps its isotopic envelope / mobility spread rather than being clipped at a hard polygon boundary; pass mz_pad=0.0, im_pad=0.0 for a hard cutoff. DiaMs1WindowGate (diaPASEF) — keeps only MS1 points inside the union of the isolation windows ( DiaFrameMsMsWindows ); everything outside is a precursor the method never isolates. Windows are padded in physical units ( mz_pad default 5 Da, im_pad default 0.05 1/K0). No-op on ddaPASEF. Both gates now no-op (keep everything) on non-MS1 frames rather than testing fragment peaks against the MS1 precursor region (which would empty an MS2 spectrum), so they are safe to leave in a noise=[...] list applied across frame types. build_window_intervals clamps a negative scan_lo to 0 and skips boxes lying wholly outside the scan range instead of letting a negative bound wrap around via Python indexing.
Fabbri, Renato
1 files · 3.5 MB · zipdeclared
BSC Lab is an open-source sensory stimulation platform with two integrated layers: a precision multi-engine audiovisual stimulation application, and a knowledge layer built on the SSTIM ontology (OWL class hierarchy, multilingual SKOS vocabulary, SHACL validation shapes, exposure module, and external alignments) published at https://w3id.org/sstim under CC BY 4.0. Each tagged release is archived here for citation.
Bradley, Alex
1 files · 12 MB · zipdeclared
R package for reconstructing precipitation hydrogen isotope ratios from sedimentary leaf-wax n-C29 alkane measurements using the spatially-aware hierarchical Bayesian calibration of Bradley (2026). Inverts the forward calibration with full uncertainty propagation, combining analytical error, residual variance, slope posterior, and the spatial Gaussian process intercept. Supports per-record change detection with autocorrelation-adjusted thresholds and a four-level claim taxonomy from measurement-level changes through uniquely- attributable precipitation-isotope claims.
Juniper L. Simonis
1 files · 86 MB · zipdeclared
Tools for interacting with the publicly available California Delta Fish Salvage Database, including continuous deployment of data access, analysis, and presentation.
Padilla-Villanueva, Johel
1 files · 457 KB · zipdeclared
Academy Learning Tau is an open educational and research software platform that operationalizes the Systemic Tau framework and the Discrete Extramental Clock (RECD) for complex time-series analysis. It provides ordinal metrics (τ_s, nested RECD levels), classical early-warning signals for dual reading, surrogate null models, multilingual pedagogy (Spanish, English, French), and a reproducible Streamlit laboratory for teaching and exploratory research.
Aouladali, Amal · Alaoui, Souad · Hnini, Abdelhalim
4 files · 491 KB · zipdeclared
Replication code and measured artifacts for KeystrokeAuthChain: a keystroke-dynamics continuous-authentication Transformer evaluated on the public CMU Killourhy-Maxion dataset (EER 7.47%), coupled to an on-chain audit layer that anchors each authentication decision on the Ethereum Sepolia testnet for non-repudiation and third-party-verifiable auditing. Includes the recognition pipeline, paired baselines, significance tests, the audit smart contracts and gas measurements, and the real end-to-end anchoring run (10 genuine decisions, on-chain decisionHash verified 10/10). The CMU dataset is not redistributed here (see data/README.md inside the archive).
Padilla-Villanueva, Johel
1 files · 356 KB · zipdeclared
Plataforma web educativa y de investigación para el paradigma Tau Sistémica y el Reloj Extramental Discreto (RECD). Incluye un laboratorio interactivo para el análisis de señales biológicas, ecológicas y financieras mediante ordinal patterns, EWS y TDA.
Bradley, Alex
1 files · 10 MB · zipdeclared
Posterior draws from 14 hierarchical Bayesian leaf-wax-to-precipitation calibration models, fit in Stan on the frozen calibration run c2_run_20260626 (n = 1128 observations; Bradley 2026, Communications Earth and Environment). Each file is a posterior::draws_df with 1000 stratified draws across 8 MCMC chains. Spatial models include 125 per-knot latent z values on a Fibonacci sphere lattice. Used by the leafwax R package as the full-resolution backing data; the package ships a 100-draw fixture and downloads the full posteriors from this archive on first use. Supersedes the v1.x deposit (v10 / n = 1129 lineage).
Suchanek, Eric G., PhD
1 files · 5.3 MB · zipdeclared
MemoryKG builds a hybrid semantic + structural knowledge graph from Markdown and plain-text document corpora. It chunks text semantically, discovers structural and semantic relationships between sections and chunks, stores them in SQLite, and augments retrieval with vector embeddings via LanceDB. It also supports a conversational memory layer — ingesting and indexing agent turns, consolidating them into summaries, and enabling semantic recall across sessions via MCP-based AI integration.
DHI
3 files · 28 GB · pdf, zipdeclared
MIKE 3 Flow Model FM is a 3D hydrodynamic modeling system based on a flexible mesh approach, used for oceanographic, coastal, and estuarine applications. This repository includes a model setup and 2-year model results for Øresund (the strait between Denmark and Sweden), observational data, and code for model validation. This dataset is part of the WaterBench series by DHI, supporting open research on water-related challenges. It is intended for educational and research purposes, including model validation, parameter calibration, and machine learning applications. Results should not be used for decision-making. Files: README: Description of dataset with details on citations, data processing, and background information. WaterBench-MIKE3HD-Oresund.zip : model setup, input data (e.g., boundary conditions, wind), observational data, and code for data exploration and model validation. MIKE3HD-Oresund-output.zip : 2-years model result files (~28 GB).
Gómez, Raimundo Elías · Miño, María Gabriela
1 files · 2.3 MB · zipdeclared
Data and analysis code for a population-level study of the formal organisational landscape of the Argentine province of Misiones, 1901-2025. The deposit contains the Misiones subset of the Registro Nacional de Sociedades (14,278 organisations; 14,169 dated), Argentina's annual real-GDP growth series 1961-2025 (World Bank, World Development Indicators, series NY.GDP.MKTP.KD.ZG), and the canonical model-free analysis: Shannon diversity with bootstrap confidence intervals, juridical-form composition with exact Poisson intervals, a functional disaggregation of cooperatives, and a permutation test of sub-departmental concentration. The central finding is a contingency asymmetry: self-sustaining commercial forms diffuse secularly across state projects of opposed political orientation, whilst the cooperative, whose creation depends on a continuous social-economy apparatus, surges when that apparatus is sustained and returns to its closed-channel share when it is withdrawn. That share occupies 4.1-4.8 per cent of new registrations in every period in which the channel through which cooperatives are constituted was closed, and 14.8-25.5 per cent in every period in which it was open, without overlap - whilst the same periods overlap in mean output growth, so the macro-economic cycle does not separate what the channel separates. An earlier correspondence-analysis and sequence pipeline was abandoned because entering the period of creation as an active variable makes the principal axis an axis of time and manufactures a spurious terminal convergence. That pipeline and the diagnostics establishing the defect are retained in archive/ for transparency; no result in the paper is computed from them. The ARCA fiscal-status enrichment ( enriquecido_arca_v2.parquet , tab_fiscal_status.csv ) is likewise retained but supports no claim in the paper: the 2024 regulatory action it would have to control for suspends an entity's authorisation to operate, which is a different register from the tax record the measure reads. Code is released under the MIT licence; data and figures under CC-BY-4.0.
Pandey, Devansh · Narasimhan, Vagheesh M.
1 files · 112 KB · zipdeclared
Analysis code for a five-layer human-genetic study of the lipid-coronary artery disease axis in UK Biobank (GWAS, LD-score genetic correlation, multivariable Mendelian randomization, colocalization, SuSiE fine-mapping, and exome-wide rare-variant burden), with two-sample MR replication against FinnGen. Code only; no UK Biobank or individual-level data are included.