Version 2.0.0 of the Composite Site Productivity Index (CSPI) for the 48 conterminous United States at 30 m (EPSG:4326). Where v1.0.0 provided a single-metric Environmental Site Index, v2.0.0 delivers a true multi-metric composite. The headline product (CSPI_v2_5component_30m.tif) is the z-score weighted average of five independently fitted productivity layers, sigma-clipped at +/-3 and rescaled to 0-100: (1) Environmental Site Index (ESI v7), (2) Biomass Growth Increment from FIA remeasurement pairs, (3) asymptotic aboveground biomass (Chapman-Richards), (4) Miami-model NPP from PRISM climate, and (5) MOD17 satellite NPP. Both NPP formulations are retained as independent components because each captures a distinct productivity signal and prior analysis showed both perform well. Coverage correction: an earlier internal v2 composite used only the Miami NPP, which had been derived from a truncated climate input and was therefore empty over New England and the coastal Pacific Northwest. The climate inputs were regenerated from intact source data and MOD17 NPP was added as a second NPP component; this release has complete CONUS coverage including northern Maine, Oregon, and Washington. Forest-masked, 1 km aggregate, and composite-scale uncertainty layers accompany the headline composite, along with the ESI predicted-site-index products and the component input layers.
Lemor, Antoine · Pillod, Alizée · Taylor, Matthew · et al.
266,271 rows × 10 cols
7 numeric · 3 text
The Canadian Climate Framing (CCF) Database is a comprehensive, machine-learning-annotated corpus of climate-change media coverage in Canada. It comprises 266,271 articles from 20 major Canadian newspapers (1978-2024) processed into 9,198,958 two-sentence analytical units (82.9% English, 17.1% French). Each unit is annotated across 65 hierarchical categories by 128 BERT and CamemBERT classifiers, with a macro F1 of 0.866 on a 1,000-sentence gold standard double-coded by an independent annotator (Gwet's AC1 = 0.894, Krippendorff's α = 0.698, Cohen's κ = 0.596 on the 400 blind sentences). Each category receives an A/B/C reliability tier summarising annotation quality from classifier performance and inter-coder agreement. The deposit ships six relational tables (bibliographic metadata, sentence-level annotations, named-entity rollups, article-level aggregates, per-category reliability tiers, and 9,462,845 BAAI/bge-m3 sentence-and-title embeddings). Raw newspaper text is excluded for copyright reasons; bibliographic coordinates (media, date, title, author, page_number) are sufficient for any researcher with institutional access to Factiva, Eureka.cc or ProQuest Canadian Major Dailies to recover the original sentences. This deposit accompanies a methodology paper currently under revision at Scientific Data (Nature Portfolio). This deposit is the Apache Parquet mirror of the canonical PostgreSQL edition (cross-referenced in Related identifiers ). Each of the six relational tables is provided as a standalone .parquet file with ZSTD compression; the 1024-dimensional BAAI/bge-m3 embedding column is materialised as a list<float> , and JSONB entity arrays are serialised as UTF-8 JSON strings. The schemas are otherwise identical to the PostgreSQL edition. The Parquet bundle is readable natively by pandas , polars , R/ arrow , DuckDB , and Spark without any database backend: import pandas as pd agg = pd.read_parquet('CCF_article_aggregates.parquet') emb = pd.read_parquet('CCF_sentence_embeddings.parquet') The HNSW index that ships with the PostgreSQL edition is not transferable to Parquet; brute-force cosine similarity remains tractable on the embedding column (≈ 9.46 M × 1024 float16). The full annotation pipeline, training data, manual-annotation JSONL, intercoder-reliability benchmark, methodology manuscript (LaTeX sources + PDF), and reproducibility scripts are bundled with this deposit as ccf_code_and_paper.tar.gz . The same materials are also available on the project's OSF companion deposit ( 10.17605/OSF.IO/Q5W47 ) and on the development mirror at GitHub .
Title Quality-controlled gradient-tower meteorological profiles and surface-flux estimates from the Huancayo Geophysical Observatory, Peruvian central Andes Alternative short title HYGO gradient-tower surface-flux dataset Resource type Dataset Version v1.0 Creators Flores-Rojas, José Luis; Pérez Tello, María; Fashé-Raymundo, Octavio; Pareja Quispe, David; Eche Llenque, José Carlos; Silva, Yamina; Zuñiga Huaman, Gerson Description This dataset contains quality-controlled gradient-tower meteorological profiles and surface-flux estimates from the Huancayo Geophysical Observatory (HYGO) of the Geophysical Institute of Peru (IGP), located in the Mantaro Valley of the Peruvian central Andes. HYGO is a high-altitude agricultural and atmospheric observatory representative of complex Andean terrain, strong diurnal forcing, seasonal moisture contrasts, and mountain-valley circulations. The dataset was developed to support reproducible analysis of near-surface atmospheric structure and turbulent exchange in complex terrain. Native 1-min observations of air temperature, relative humidity, wind speed, and wind direction were processed through a documented workflow that includes timestamp auditing, primary meteorological quality control, conservative bit-mask flagging, thermodynamic derivation, 30-min aggregation, flux-specific pre-calculation quality control, dual-method turbulent-flux estimation, method-status diagnostics, and post-calculation plausibility filtering. The released products include cleaned 1-min tower observations, per-sample QC flags, derived thermodynamic variables, 30-min aggregated profiles, and surface-flux estimates obtained with two aerodynamic approaches: Monin-Obukhov Similarity Theory (MOST) and an anchored multi-layer Bulk Richardson Number method (BRN_ANC). The flux products include friction velocity, sensible heat flux, latent heat flux, Obukhov length, bulk Richardson number, method-status flags, post-calculation QC flags, raw method outputs, and post-QC-filtered outputs. This structure allows users to distinguish between input-profile limitations, numerical method failures, and physically implausible flux estimates. The gradient-tower system includes measurements at multiple levels between 2 and 29 m above ground level. For the flux-gradient calculations, the 2, 6, 12, and 24 m levels were used to construct vertical profiles of wind speed, temperature, humidity, and virtual potential temperature. The 18 and 29 m levels were used for wind-direction information where available but were not included in the main flux-gradient calculations. The dataset is intended for boundary-layer research, land-atmosphere interaction studies, evaluation of surface-layer parameterizations, comparison of MOST and Richardson-number methods, model validation, agricultural micrometeorology, frost-risk assessment, drought-related studies, and development of reproducible workflows for high-frequency meteorological tower data. Dataset period [Insert final period, e.g., 15 May 2018 to 30 April 2026] Geographic coverage Huancayo Geophysical Observatory, Mantaro Valley, central Peruvian Andes Latitude: [insert final latitude, e.g., -12.04145] Longitude: [insert final longitude, e.g., -75.31875] Elevation: [insert final elevation, e.g., 3315 m a.s.l.] Temporal resolution 1 min for native and cleaned meteorological observations. 30 min for aggregated profiles and turbulent-flux products. Main variables Air temperature Relative humidity Wind speed Wind direction Atmospheric pressure Saturation vapour pressure Actual vapour pressure Water-vapour mixing ratio Specific humidity Virtual potential temperature Friction velocity Sensible heat flux Latent heat flux Obukhov length Bulk Richardson number Input QC flags Method-status flags Post-calculation QC flags 30-min availability diagnostics Processing summary Raw 1-min tower observations were time-sorted, audited for duplicate timestamps, and regularized to a 1-min temporal grid when required. A primary meteorological QC system generated per-sample bit-mask flags for missing values, range violations, step changes, persistence, spikes, calm wind, humidity inconsistency, and resample-inserted timestamps. Hard-fail values were removed under a conservative rule: RANGE or simultaneous STEP and SPIKE. Contextual flags were retained for diagnostic use. Thermodynamic variables were derived after QC, including vapour-pressure variables, specific humidity, and virtual potential temperature. Cleaned 1-min profiles were aggregated to 30-min profiles with availability diagnostics. Flux-specific pre-calculation QC screened each 30-min profile before flux estimation. MOST and BRN_ANC flux estimates were computed independently from the same eligible profiles. Method-status flags recorded numerical success, non-convergence, invalid profile slopes, Richardson-number exceedance, and other execution outcomes. Post-calculation QC retained physically plausible flux estimates and masked non-passing values in the final filtered output columns. Raw method outputs were preserved separately to support diagnostic audits and sensitivity analyses. File contents [Edit this list to match the final Zenodo upload.] cleaned_1min_tower_data.[nc/csv] Cleaned 1-min meteorological observations and primary QC flags. derived_thermodynamic_variables.[nc/csv] Pressure, vapour-pressure variables, mixing ratio, specific humidity, and virtual potential temperature. aggregated_30min_profiles.[nc/csv] Thirty-minute mean profiles and data-availability diagnostics. surface_fluxes_MOST_BRN_ANC_30min.[nc/csv] MOST and BRN_ANC flux estimates, method-status flags, post-QC flags, raw outputs, and filtered outputs. qc_flag_dictionary.[csv/json] Definitions of primary QC bit masks, pre-calculation QC flags, post-calculation QC flags, and method-status flags. processing_scripts.[zip] Python scripts used for QC, thermodynamic derivation, aggregation, MOST, BRN_ANC, post-QC, diagnostics, and figures. environment.[yml/txt] Software environment and package dependencies required to reproduce the workflow. README.md Dataset description, file structure, variable names, units, QC interpretation, and recommended use. Recommended citation Flores-Rojas, J. L., Pérez Tello, M., Fashé-Raymundo, O., Pareja Quispe, D., Eche Llenque, L. Suárez Salas, J. C., Silva, Y., and Zuñiga Huaman, G. ([year]). Quality-controlled gradient-tower meteorological profiles and surface-flux estimates from the Huancayo Geophysical Observatory, Peruvian central Andes (Version v1.0) Keywords gradient tower; surface energy fluxes; quality control; Monin-Obukhov Similarity Theory; MOST; Bulk Richardson number; BRN_ANC; atmospheric surface layer; boundary layer; turbulent fluxes; sensible heat flux; latent heat flux; friction velocity; tropical Andes; Mantaro Valley; Huancayo Geophysical Observatory; HYGO; Peru; micrometeorology; land-atmosphere interactions; reproducible workflow License [Recommended: Creative Commons Attribution 4.0 International, CC BY 4.0, if allowed by your institution and funder.] Related identifiers Is supplement to: [insert article DOI after publication] Is documented by: [insert manuscript/preprint DOI if available] Is supplemented by: [insert software DOI if scripts are archived separately] Is version of: [insert previous Zenodo DOI if this is an updated version] Funding Instituto Geofísico del Perú; PROCIENCIA project "Fortalecimiento del Laboratorio de Microfísica Atmosférica y Radiación para el estudio de la interacción superficie-atmósfera en una zona agrícola de los Andes Centrales del Perú, en el contexto de cambio climático" (LAMAR), Contract No. PE501086050-2023-PROCIENCIA-BM. Notes Users should treat the flux estimates as gradient-based products, not as direct eddy-covariance measurements. MOST and BRN_ANC estimates are provided together to support method comparison and uncertainty assessment. Strongly stable, weak-wind, transition-period, and horizontally heterogeneous conditions may increase uncertainty. Users are encouraged to use the QC flags, method-status flags, and raw-output variables when performing sensitivity analyses or applying stricter filters.
Abstract Background : Suspended particulate matter smaller than 10 μm (PM10) contained in atmospheric aerosols is a cause for concern as it can reach the human respiratory system and potentially have adverse effects on human health. In particular, particles smaller than 2.5 μm ( PM2.5 ) contain many combustion-derived substances that generate free radicals harmful to the human cardiopulmonary system. In recent years, the increase in combustion particles due to the use of fossil fuels has become a global problem. The author examined meteorological factors related to the dynamics of PM10 and PM2.5 in Tokyo, Japan's largest city. Method : The daily average values for PM10 and PM2.5 at the Tokyo National Environmental Observatory from April to November 2019 were obtained from the Tokyo Metropolitan Government Bureau of Environment. The daily data such as ambient temperature, relative humidity, wind speed and precipitation of the Tokyo Meteorological Observatory was downloaded from the Japan Meteorological Agency. Result and Discussion : PM10 concentration showed a significant positive correlation with ambient temperature and relative humidity, and a weak negative correlation with precipitation and wind speed. PM2.5 concentration showed a significant positive correlation with ambient temperature, a weak negative correlation with relative humidity, and a significant negative correlation with precipitation and wind speed. The results of the stepwise linear regression analysis revealed that for PM10, temperature and relative humidity were significant positive independent variables, while for PM2.5, temperature was a significant positive independent variable, and relative humidity and wind speed were significant negative independent variables. The daily average concentration trends of particulate matter from April to November, high concentrations were frequently observed in July and August, when temperatures are high. This is presumed to be the reason why particulate matter concentrations showed a significant positive correlation with temperature. High concentrations of particulate matter during periods of high temperatures may be related to increased electricity consumption and the urban heat island effect. The positive correlation observed between relative humidity and PM10 concentration may suggest the formation of secondary inorganic aerosols. On the other hand, the negative correlation observed between relative humidity and PM2.5 concentration suggests the influence of rain-out and wash-out due to rainfall. The negative correlation between wind speed and particulate matter concentration was thought to represent ventilation. Conclusion: The author investigated the relationship between various meteorological factors and particulate matter concentrations in the metropolitan area of Tokyo, and also considered the dynamics of particulate matter related to the recent occurrence of the urban heat island phenomenon. Keywords : urban aerosols; PM10; PM2.5; ambient temperature; relative humidity; wind speed; precipitation; heat island phenomenon
Vivek Amuthan S · K L Prakash · Sharanya S V · et al.
1 files · 812 KB · pdfdeclared
Rapid urbanisation in tropical megacities has engendered complex atmospheric regimes characterised by extreme spatiotemporal heterogeneity. This study presents a comprehensive statistical characterisation of the ambient air quality matrix in Bengaluru, India (2018-2024), leveraging a high-density network of 13 Continuous Ambient Air Quality Monitoring Stations (CAAQMS). Moving beyond aggregate compliance metrics, this study employs an integrated statistical framework using the Mann-Kendall trend test, Coefficient of Divergence (COD), and Multi-Dimensional Scaling (MDS). Results indicate a 'fractured airshed' (defined as a domain exhibiting high spatial heterogeneity, COD > 0.20): the industrial node (Peenya) recorded a statistically significant decline in PM2.5 (Sen's Slope: -2.75 μg/m³/year), contrasting with a statistically significant increased trend of Ammonia (NH₃) in residential zones (+1.44 μg/m³/year). Spatial analysis revealed a high Coefficient of Divergence (COD = 0.42) between industrial and residential zones, refuting the hypothesis of a homogenous urban airshed. Further, source diagnostics quantified a weekend effect with an 18.4% reduction in NO₂ (p < 0.01), isolating the vehicular contribution to the pollution burden. These findings underscore the critical need for a Zonal Air Quality Management (ZAQM) framework.
Aimed at year-round recording of the chemical aerosol composition in central Antarctica, an unattended operating aerosol sampler was successfully deployed at the EPICA deep drilling site in Dronning Maud Land (Kohnen Station). Analyses of teflon/nylon filter packs consecutively collected over bi-weekly intervals during the February 2003 to December 2005 period allowed to evaluate seasonal concentration variations of methane sulphonate (MS), Cl-, NO3-, non-sea salt (nss-)SO4**2- and Na+, while NH4+ and mineral dust related ion results remained below detection limits. For MS and nss-SO4**2 distinct late summer maxima around 44 and 200 ng/m**3, respectively, were found, while (total) NO3- showed a broad November maximum of about 52 ng m**-3. In contrast, the highest concentrations of Na+ with peak values of up to 160 ng/m**3 were observed during the winter half year. The seasonality of these species broadly coincided with long-term observations at the coastal Neumayer Station, including surprisingly comparable NO3- levels. However, the biogenic sulphur and sea salt concentrations were lower at Kohnen by typically a factor of 2-3 and 10, respectively. The arrival of sea ice derived sea salt particles at Kohnen could not clearly detected, since even during mid-winter the nss-SO4**2- to Na+ ratio was generally too high to unambiguously identify a sulphur depleted sea salt SO4**2- fraction.
Aerosol samples collected over the North Atlantic from ship were analysed for Sodium, Magnesium, Potassium, Calcium and Chloride. A found dependence of sea salt concentrations from wind velocity is compared with earlier results. The mean of the ratio Cl/Na was close to that for sea water; the Mg-, K- and Ca-concentrations in the aerosol, however, were enriched with respect to sea water. It is shown that continental advection influences the measured aerosol components over the North Atlantic.
Lemor, Antoine · Pillod, Alizée · Taylor, Matthew · et al.
2.2 MB
The Canadian Climate Framing (CCF) Database is a comprehensive, machine-learning-annotated corpus of climate-change media coverage in Canada. It comprises 266,271 articles from 20 major Canadian newspapers (1978-2024) processed into 9,198,958 two-sentence analytical units (82.9% English, 17.1% French). Each unit is annotated across 65 hierarchical categories by 128 BERT and CamemBERT classifiers, with a macro F1 of 0.866 on a 1,000-sentence gold standard double-coded by an independent annotator (Gwet's AC1 = 0.894, Krippendorff's α = 0.698, Cohen's κ = 0.596 on the 400 blind sentences). Each category receives an A/B/C reliability tier summarising annotation quality from classifier performance and inter-coder agreement. The deposit ships six relational tables (bibliographic metadata, sentence-level annotations, named-entity rollups, article-level aggregates, per-category reliability tiers, and 9,462,845 BAAI/bge-m3 sentence-and-title embeddings). Raw newspaper text is excluded for copyright reasons; bibliographic coordinates (media, date, title, author, page_number) are sufficient for any researcher with institutional access to Factiva, Eureka.cc or ProQuest Canadian Major Dailies to recover the original sentences. This deposit accompanies a methodology paper currently under revision at Scientific Data (Nature Portfolio). This deposit is the canonical PostgreSQL edition. It contains a pg_dump -Fd directory archive (compressed into a single .tar file) of the six relational tables, including the pgvector extension and HNSW cosine indexes for sub-second semantic-similarity search. Restoration is a one-liner: tar -xf CCF_Database.tar && createdb CCF_Database && psql -d CCF_Database -c 'CREATE EXTENSION IF NOT EXISTS vector;' && pg_restore -d CCF_Database --no-owner --no-privileges -j 8 CCF_Database_dump A column-oriented Apache Parquet mirror of the same six tables is available as the sister deposit on Zenodo (cross-referenced in Related identifiers ). The Parquet mirror is recommended for users without PostgreSQL access (it is directly readable by pandas, polars, R/arrow, DuckDB, and Spark). The full annotation pipeline, training data, manual-annotation JSONL, intercoder-reliability benchmark, methodology manuscript (LaTeX sources + PDF), and reproducibility scripts are bundled with this deposit as ccf_code_and_paper.tar.gz . The same materials are also available on the project's OSF companion deposit ( 10.17605/OSF.IO/Q5W47 ) and on the development mirror at GitHub . Requirements: PostgreSQL 16 or 17 with pgvector ≥ 0.8.2 (for halfvec(1024) storage of the sentence embeddings).
During the 1965 Atlantic Expedition of the "Meteor“ concentrations of various atmospheric trace gases were measured. The following gases were considered: carbon dioxide (CO2), sulfur dioxide (SO2), nitrogene dioxide (NO2), and nitric oxide (NO). The air whereof these components were measured was sucked in from a height of 14 m above the surface of the sea. The results allow conclusions upon the long term global increase of the atmospheric CO2 content, the meridional distribution of the CO2 on the Atlantic Ocean, and the dependance of its concentration upon the time of the day and the thermal structure of the atmosphere. Attempts at determining concentrations of sulfur dioxide and nitric oxide of non-continental origin failed at large. Concentrations of NO2, however, could succesfully be measured.
Global GEOS-Chem simulations at 2° × 2.5° resolution were used to generate O3 concentrations constrained by OMI NO2 observations from the NASA standard product for 2006–2016. The ozone simulations were driven by NOx emissions from https://doi.org/10.7910/DVN/HVT1FO
GLC_FCS60 is the first global 60-m land cover product with a fine classification system developed using comprehensive change detection. It employs a refined classification system inherited from the GLC_FCS30D product, which contains 35 land-cover classes and covers the years 1975 and 1980. Specifically, it is developed by combining an improved comprehensive change detection method, local adaptive classification models, and a series of post-classification processing methods, primarily using Landsat MSS time-series imagery. Validation results show that the 1975 product achieves overall accuracies of 80.13% for the 10 basic classes and 71.60% for the 17 level-1 classes, whereas the 1980 product achieves corresponding accuracies of 80.08% and 70.71%, respectively. The GLC_FCS60 dataset has been stored by a total of 961 independent 5°×5° geographical tiles, and the tile names as " GLC_FCS60_19751980_E *** N## ", in which the "***" and "##" illustrate the longitude and latitude coordinates of the upper left corner of the tile data.
Piel, Claudia · Weller, Rolf · Huke, Michael · et al.
98 rows × 14 cols · 13 KB
8 text · 5 numeric · 1 datetime
During three summer campaigns in January/February 2000, 2001, and 2002 the ionic composition of the aerosol at the European Project for Ice Coring in Antarctica (EPICA) deep-drilling site at Kohnen Station was measured in daily resolution. In 2000 and 2002 we observed mean (±std) non-sea-salt sulfate (nss-[SO4]2-) concentrations of 353 ± 100 ng/m**3 and 320 ± 250 ng/m**3, as well as methane sulfonate (MS) concentrations of 59 ± 36 ng/m**3 and 74 ± 80 ng/m**3, respectively. For the summer campaign in 2001, significantly lower nss-[SO4]2- and MS levels of 164 ± 150 ng/m**3 and 19 ± 12 ng/m**3, respectively, were typical. The mean MS/nss-[SO4]2- ratio ranged from about 0.1 to 0.2. MS and nss-[SO4]2- concentrations and their variability were roughly comparable to coastal stations at summer. Supported by air mass back trajectory analyses, this finding documented an efficient long-range transport to Kohnen via the free troposphere. MS/nss-[SO4]2- ratios exhibited a strong dependence on the MS concentration with systematically higher ratios at higher MS concentrations, a peculiarity which is also evident in a firn core drilled at this site.
This dataset quantifies the uncertainty in mapping late-successional and old-growth (LSOG) forest across the approximately 4.2 million hectares of Maine's unorganized townships, and tests whether LSOG is rapidly disappearing. Three to four independent, credible mapping methods are compared on a common 100 m grid: (M1) a reproduction of the Hagan et al. (2026) airborne-LiDAR canopy random forest, rebuilt from their public Zenodo deposit; (M2) a logistic model of the FIA field-structure LSOG class on Potapov (GEDI-calibrated) canopy height; (M3) a direct canopy-height threshold; and (M4) the FIA structural class imputed to every pixel via USFS TreeMap (2016, 2020, 2022). Version 1.2.0 additions. This version adds the materials behind the formal Ecosphere Comment on Hagan et al. (2026): (a) a cross-validated accuracy assessment (AUC) of each mapping approach on the original authors' own training plots, showing that high training accuracy does not transfer to agreement among independent maps; (b) an FIA design-based estimate of older forest with sampling-error confidence intervals, the unbiased ground reference the original analysis lacked, putting older forest at about 3.9 percent (3.3 to 4.6) and rising, including on private commercial timberland; (c) a threshold-sensitivity sweep and a 20-seed reproduction ensemble; (d) an ownership-resolved breakdown (private commercial versus public); (e) a hex-scale (8 km) summary of cross-method disagreement; and (f) the Comment manuscript and Supporting Information. Headline findings. Credible methods disagree by roughly 2.8 times on how much LSOG exists and on the location of most LSOG hectares, while agreeing closely on the rare, well-defined old-growth core. Protecting the top 5 to 20 percent of hectares by one map versus another overlaps on only 16 to 30 percent of the ground, so single-map patch-level prioritization for large expenditures is fragile. The design-based FIA estimate and TreeMap imputation both show older forest stable to increasing rather than rapidly declining; the apparent loss reported elsewhere is a gross harvest flux, not a net stock decline. Contents. Derived 100 m GeoTIFFs (reproduced Hagan class, v5.1-GEDI probability, TreeMap class, a per-cell method-consensus layer), summary tables (area by method, pairwise agreement, concordance, prioritization fragility, AUC by approach, design-based older-forest trend with CIs, ownership breakdown, and FIA validation), the analysis R scripts, quick-look figures, the Ecosphere Comment manuscript and Supporting Information, and a full methods-and-findings report (PDF). Privacy. No FIA plot coordinates are included; all products are derived rasters or aggregate summary tables. Caveats: the robust temporal signal is direction rather than precise rate; FIA stand age is modeled, so a structural large-tree domain is reported alongside the age domain; cross-validated intervals are best read as lower bounds because plots are spatially dispersed but not independent. See the README and report for full methods, provenance, and limitations. Version 1.12.0 additions. The cross-map comparison is refined to independent remote-sensing operationalizations only. (a) A three-map remote-sensing ensemble over Maine on a common 100 m grid: reproduced Hagan airborne-LiDAR (any-LSOG 21.9 percent), an FIA-structure class on Potapov GEDI-calibrated spaceborne canopy height (14.0 percent), and the ORNL/Bruening national old-growth stratum (36.1 percent); the three span a 2.6-fold range and agree on only 2.7 percent of flagged hectares, with the ORNL stratum spatially uncorrelated with the structure maps. The USFS TreeMap imputation is reclassified as a second FIA-anchored accounting, reported with the design-based estimate rather than as an independent map. (b) A design-based estimate of LSOG itself: integrated any-LSOG 14.1 percent (12.9 to 15.3) and strict four-axis true LSOG 3.1 percent (2.5 to 3.7) of Maine forestland, the airborne map exceeding even the inclusive ground estimate. (c) A balanced-model LSOG probability surface at 100 m, with a binary class calibrated to the design-based area to bound over-prediction. (d) A multi-objective support vector regression pilot tracing the Pareto front of total versus systematic (attenuation) error. Derived rasters, the three-map agreement layer, the ORNL stratum reprojected to the study grid, tables, R scripts, the updated Comment, and the companion manuscript are included. No FIA plot coordinates are included. Version 1.13.0. Final consolidated release. Adds: a rare-class remedy menu for the reproduced random forest (default vs class weighting vs balanced sub-sampling vs voting-threshold; old-growth detection 0.24 to 0.82, mapped old-growth area 1.0 to 2.3 percent); the definitive five-model LSOG probability map for Maine with across-model uncertainty and reference reserves (MNAP/TNC network, Baxter, Big Reed) over real state and county boundaries; design-based 95 percent confidence intervals for forest-type and ecoregion representation and for disturbance shares; a full robustness/stress-test matrix; and the copy-edited, sole-authored Comment, companion manuscript, and Maine Forest Products Council technical report. Authorship updated to Aaron R. Weiskittel.
This public analysis script summarizes the core analytical workflow used in the manuscript. Using a time-stratified case-crossover design, it fits distributed lag non-linear models with conditional logistic regression to estimate cumulative associations between day-to-day changes in actual vapor pressure, absolute humidity, and relative humidity and pediatric allergic rhinitis visits. The script further implements subgroup analyses for actual vapor pressure change by sex, age, season, and background absolute humidity, as well as sensitivity analyses for lag windows, covariate adjustment schemes, and linear pollutant terms. It also summarizes odds ratios and attributable fractions for preprocessed future scenario Δeₐ series to reproduce the main epidemiological and projection-based results.
Version 3.3 update (June 2026): This version extends LCSPP-MODIS from the original 2001-01-2023-12 record through 2026-05-31 (new observations 2024-01-01 to 2026-05-31; biweeks 202401a-202605b) and refines the global land mask. The v3.2 baseline pipeline (Fang et al. 2025) is unchanged in principle: MCD43C1.061 BRDF → biweekly max-NDVI composite (MVC) normalized to SZA=45° → snow masking → HANTS gap-fill → LCREF (red/NIR reflectance) → SIF-trained neural network → LCSPP (clear-sky instantaneous, clear-sky daily, all-sky daily). Updates in v3.3: Temporal extension : new MODIS observations (MCD43C1.061) for 2024-01 through 2026-05. The historical 2001-2023 MVC was reconstructed from the published v3.2 LCREF (observed, QA=0 pixels), since the original raw biweekly composites were no longer archived; the reconstruction reproduces the published record at observed pixels (see validation below). Gap-filling : HANTS is now applied over a 9-year moving window (±4 years, 216 biweeks) rather than a single full-series fit, better capturing low-frequency variability and accommodating the open-ended record. Windows are best-effort centered and clamped to complete years at the record end; the incomplete 2026 is isolated to one shifted window that preserves full annual periodicity. The QA system is unchanged: lcspp_qa = max(red_qa, nir_qa), with 0 = observation, 1 = high-quality HANTS gap-fill, 2 = lower-quality (mean-seasonal-cycle) climatology fill, 3 = no data. Snow ancillary layer : a snow_masked layer (1 = snow-contaminated pixel masked before gap-fill) is provided for the 2024-onward biweeks. (set to 0 for the 2001-2023 period) All-sky daily / ERA5 SSRD : the all-sky daily rescaling uses ERA5 surface solar radiation downwards (SSRD), rebuilt from CDS daily-statistics (2001-2017) and the ARCO-ERA5 archive (2018-2026). Final ERA5 (expver 1) is confirmed through 2026-03-31; 2026-04 onward uses preliminary ERA5T, flagged per file in the era5_ssrd_status attribute and refreshable once final ERA5 is published. Validation (v3.3 vs v3.2, 2001-2023 overlap) : observed (QA=0) pixels reproduce v3.2 essentially exactly (MAE ≈ 2×10 -8 ); pixelwise spatial correlation r ≈ 0.998; growing-season area-weighted annual means and 8-ecosystem 1°×1° box-mean trajectories agree at r = 1.00; all-sky daily (the SSRD path) differs by ≈ 0 (r = 1.00). Differences are confined to gap-filled pixels, plus an expected increase in trailing-window climatology (QA=2) coverage during 2020-2023. The original version 3.2 description follows below. Paper Reference: Fang, J., Lian, X., Ryu, Y., Jeong, S., Jiang, C., & Gentine, P. (2025). A long-term reconstruction of a global photosynthesis proxy over 1982-2023. Scientific data , 12 (1), 372. https://doi.org/10.1038/s41597-025-04686-6 Usage Notes : This is the updated LCSPP dataset (v3.2), reconstructed using the MODIS record from 2001-2023. Previously referred to as "LCSIF," the dataset was renamed to emphasize its role as a SIF-informed long-term photosynthesis proxy derived from surface reflectance and to avoid confusion with directly measured SIF signals. The MODIS-based LCSPP is generated as an ancillary product to complement and benchmark the LCSPP-AVHRR product from 1982-2023. Key updates in version 3.2 include: Improved Calibration : Enhanced consistency in calibration methods, addressing technical limitations in version 3.1 including applying more stringent quality filtering and snow masks. Quality Flags : New quality flag layer enables users to identify whether a pixel is derived from observed surface reflectance (QA=0), high-quality gap-filled values (QA=1), lower-quality gap-filled based on the mean seasonal cycle (QA=2), or missing entirely (QA=3). We advice the user to rely only on observed and high-quality gap-filled values for their analyses. Extension to include observations from the year of 2023. LCSPP-AVHRR repositories can be accessed via the following links: LCSPP-AVHRR v3.2 (1982-2000): 10.5281/zenodo.7916850 LCSPP-AVHRR v3.2 (2001-2023): 10.5281/zenodo.11906675 The user can choose between LCSPP-AVHRR and LCSPP-MODIS for the overlapping period from 2001-2023. The two datasets are generally consistent during this overlapping period, although LCSPP-MODIS shows a stronger greening trend between 2001-2023. For studies exploring the long-term vegetation dynamics, the user can either use only LCSPP-AVHRR or use a blend dataset of LCSPP-AVHRR and LCSPP-MODIS as a sensitivity test. In addition, the updated long-term continuous reflectance datasets (LCREF), used for the production of LCSPP, can be accessed using the following links: LCREF-AVHRR v3.1 (1982-2023): 10.5281/zenodo.11905959 LCREF-MODIS v3.1 (2001-2023): 10.5281/zenodo.11657458 A paper describing the technical details is available at https://doi.org/10.1038/s41597-025-04686-6 , while detailed the uses and limitations of the dataset. In particular, we note that LCSPP is a reconstruction of SIF-informed photosynthesis proxy and should not be treated as SIF measurements . Although LCSPP has demonstrated skill in tracking the dynamics of GPP and PAR absorbed by canopy chlorophyll (APARchl), it is not suitable for estimating fluorescence quantum yield. All data outputs from this study are available at 0.05° spatial resolution and biweekly temporal resolution in NetCDF format. Each month is divided into two files, with the first file "a" representative of the 1 st day to the 15 th day of a month, and the second file "b" representative of the 16 th day to the last day of a month. Abstract: Satellite-observed solar-induced chlorophyll fluorescence (SIF) is a powerful proxy for the photosynthetic characteristics of terrestrial ecosystems. Direct SIF observations are primarily limited to the recent decade, impeding their application in detecting long-term dynamics of ecosystem function. In this study, we leverage two surface reflectance bands available both from Advanced Very High-Resolution Radiometer (AVHRR, 1982-2023) and MODerate-resolution Imaging Spectroradiometer (MODIS, 2001-2023). Importantly, we calibrate and orbit-correct the AVHRR bands against their MODIS counterparts during their overlapping period. Using the long-term bias-corrected reflectance data from AVHRR and MODIS, a neural network is trained to produce a Long-term Continuous SIF-informed Photosynthesis Proxy (LCSPP) by emulating Orbiting Carbon Observatory-2 SIF, mapping it globally over the 1982-2023 period. Compared with previous SIF-informed photosynthesis proxies, LCSPP has similar skill but can be advantageously extended to the AVHRR period. Further comparison with three widely used vegetation indices (NDVI, kNDVI, NIRv) shows a higher or comparable correlation of LCSPP with satellite SIF and site-level GPP estimates across vegetation types, ensuring a greater capacity for representing long-term photosynthetic activity.
This dataset contains the estimates of potential particle emissons (CH4, CO, CO2, NH3,NOx, PM2.5, PM10 and SO2). Estimates were produced from wildfire behavior simulations.
Version 3.3 update (June 2026): This version extends LCREF-MODIS (BRDF-normalized red and near-infrared reflectance) from the original 2001-01-2023-12 record through 2026-05-31 (new observations 2024-01-01 to 2026-05-31; biweeks 202401a-202605b) and refines the global land mask. The v3.2 baseline production (Fang et al. 2025) is unchanged: MCD43C1.061 BRDF → biweekly max-NDVI composite normalized to SZA=45° → snow masking → HANTS gap-fill → red/NIR reflectance. Updates in v3.3: Temporal extension : new MODIS observations (MCD43C1.061) for 2024-01 through 2026-05. Gap-filling : HANTS is now applied over a 9-year moving window (±4 years, 216 biweeks) instead of a single full-series fit, better capturing low-frequency variability and accommodating the open-ended record. Windows are best-effort centered and clamped to complete years at the record end; the incomplete 2026 is isolated to one shifted window preserving full annual periodicity. The QA layers (red_qa, nir_qa) are unchanged: 0 = observation, 1 = high-quality HANTS gap-fill, 2 = climatology fill, 3 = no data. Snow ancillary layer : a snow_masked layer (1 = snow-contaminated pixel masked before gap-fill) is provided for the 2024-onward biweeks (0 over the 2001-2023 record). Land mask : all pixels are clamped to the MODIS IGBP land mask (MCD12C1, unioned with the LCREF-v3.2 land footprint), removing ocean/coastal bleed. Validation (v3.3 vs v3.2, 2001-2023 overlap) : observed (QA=0) pixels reproduce v3.2 essentially exactly; pixelwise spatial correlation r ≈ 0.998. Differences are confined to gap-filled pixels, plus an expected increase in trailing-window climatology (QA=2) coverage during 2020-2023. The original version 3.2 description follows below. Paper Reference: Fang, J., Lian, X., Ryu, Y., Jeong, S., Jiang, C., & Gentine, P. (2025). A long-term reconstruction of a global photosynthesis proxy over 1982-2023. Scientific data , 12 (1), 372. https://doi.org/10.1038/s41597-025-04686-6 Usage Notes : This is the updated LCREF-MODIS dataset (v3.2) consists of BRDF-normalized MODIS red and near-infrared surface reflectance. The LCREF-MODIS product was used to calibrate and benchmark the AVHRR surface reflectance to produce a temporally consistent record of surface reflectance prior to the MODIS era. It was also used to generate LCSPP-MODIS (previously known as LCSIF-MODIS) as a benchmark. Key updates in version 3.2 include: Quality Flags : New quality flag layer enables users to identify whether a pixel is derived from observed surface reflectance (QA=0), high-quality gap-filled values (QA=1), lower-quality gap-filled based on the mean seasonal cycle (QA=2), or missing entirely (QA=3). We advice the user to rely only on observed and high-quality gap-filled values for their analyses. Extension: to include observations from the year of 2023. Snow mask: we note that all pixels marked with percent_snow >0 in the original MCD43C1.v061 have been removed. This conservative approach was applied to reduce bias during cross-calibration, since unlike MODIS, AVHRR does not have a reliable snow detection algorithm. Therefore, surface reflectance values in high latitude regions are almost entirely gap-filled and should never be used for analysis for both LCREF-AVHRR and LCREF-MODIS. We encourage users to use only QA=0 and QA=1 pixels for their analysis. Alternatively, users can use LCREF-MODIS from the previous version for high latitude regions (v3.1), which did not mask out snow-covered pixles. The user can choose between LCREF-AVHRR and LCREF-MODIS for the overlapping period from 2001-2023. The two datasets are generally consistent during this overlapping period, although LCREF-MODIS shows a stronger greening trend between 2001-2023. For studies exploring the long-term vegetation dynamics, the user can either use only LCREF-AVHRR or use a blend dataset of LCREF-AVHRR and LCREF-MODIS as a sensitivity test. The LCREF-AVHRR v3.2 (1982-2023) is available at 10.5281/zenodo.11905959 The LCREF-AVHRR dataset was used as the input to generate LCSPP-AVHRR (previously known as LCSIF-AVHRR), and it can also be used to derive temporally consistant records of NDVI, NIRv, kNDVI, and other vegetation indices based on red and NIR surface reflectance variables. The user can access LCSPP products at: LCSPP-AVHRR v3.2 (1982-2000): 10.5281/zenodo.7916850 LCSPP-AVHRR v3.2 (2001-2023): 10.5281/zenodo.11906675 LCSPP-MODIS v3.2(2001-2023): 10.5281/zenodo.11657458 A paper describing the technical details is available at https://doi.org/10.1038/s41597-025-04686-6 , which detailed the uses and limitations of the dataset. All data outputs from this study are available at 0.05° spatial resolution and biweekly temporal resolution in NetCDF format. Each month is divided into two files, with the first file "a" representative of the 1 st day to the 15 th day of a month, and the second file "b" representative of the 16 th day to the last day of a month. Abstract: Satellite-observed solar-induced chlorophyll fluorescence (SIF) is a powerful proxy for the photosynthetic characteristics of terrestrial ecosystems. Direct SIF observations are primarily limited to the recent decade, impeding their application in detecting long-term dynamics of ecosystem function. In this study, we leverage two surface reflectance bands available both from Advanced Very High-Resolution Radiometer (AVHRR, 1982-2023) and MODerate-resolution Imaging Spectroradiometer (MODIS, 2001-2023). Importantly, we calibrate and orbit-correct the AVHRR bands against their MODIS counterparts during their overlapping period. Using the long-term bias-corrected reflectance data from AVHRR and MODIS, a neural network is trained to produce a Long-term Continuous SIF-informed Photosynthesis Proxy (LCSPP) by emulating Orbiting Carbon Observatory-2 SIF, mapping it globally over the 1982-2023 period. Compared with previous SIF-informed photosynthesis proxies, LCSPP has similar skill but can be advantageously extended to the AVHRR period. Further comparison with three widely used vegetation indices (NDVI, kNDVI, NIRv) shows a higher or comparable correlation of LCSPP with satellite SIF and site-level GPP estimates across vegetation types, ensuring a greater capacity for representing long-term photosynthetic activity.