The round trip
You are a dataset. Follow yourself down. Rise back up as an answer.
You are a 4.2 GB climate file sitting in a repository. Nobody has opened you. A question is on its way down to find you, and scrolling this page is the whole round trip: the descent that measures you, then the climb back up as an answer.
you are here: 4.2 GB, unopened, top of the round trip
L0 declared
Where you live now
Every catalog indexes what a publisher declared: a title, a DOI, a size, a license. Google Dataset Search, DataCite and Hugging Face all stop exactly here. Everything below the line is the bytes nobody opened, the unsolved last mile.
L0 to L1 · the deep probe
The touch: a few KB, then discard
The probe reaches into your multi-gigabyte body and reads only a header range and a footer range, kilobytes in all. A Parquet footer alone yields an exact row count and per-column min, max and null counts: kilobytes read from a multi-gigabyte body, orders of magnitude less data moved. It measures, records exactly how much it touched, then discards the bytes. Metadata-only is a custody rule, not an epistemology: no byte custody, but knowledge earned from bytes.
the ladder
Understanding is earned, rung by rung
Each level is earned from the one below, never asserted. Measured distributions and quantiles, then the role grammar, then meaning read from evidence alone with every claim cited: the hallucination firewall. Topology is a closed vocabulary of ten - tables, tensors, spatial rasters, sequences, graphs, 3D molecular structures and more - so one strategy is written once per measurable topology, never once per format.
one engine, two doors
Served byte-identically, two altitudes
Now fully measured, you are served through two front doors from one Catalog facade and one set of frozen contracts. The agent surface is the primary customer and gets a lean projection; the human Workbench gets the full record. Same truth, different altitude. The loader always runs in your runtime: Meridian is the brain, your environment is the hands.
the action layer
The right loader for what you actually are
Because the probe measured you, Meridian hands the loader for your specific topology: pandas for a tabular CSV, xarray for a NetCDF tensor, geopandas with masked nodata and real coordinate names for a raster. Not a generic snippet. And the same measured structure is served as standards on every endpoint, Croissant first.
the ML-native contract
Spoken in the format ML tools already read
Croissant is the metadata format ML datasets are converging on: an MLCommons standard, a JSON-LD layer on schema.org that tells a framework exactly what files exist, how fields are extracted, and how to load them. Hugging Face auto-generates it for every dataset, Kaggle and OpenML serve it, TensorFlow Datasets and the mlcroissant library load it, and Google Dataset Search indexes it. Meridian emits Croissant 1.1, layers in the GeoCroissant extension for geospatial data, and proves conformance in CI against the real mlcroissant library and our own SHACL shapes, with one difference no other producer has: every field statistic, array shape and spatial measurement comes from your bytes, not copied from what a publisher declared. And the same measured structure serves every standard at one endpoint: Croissant 1.1 for ML loaders, DCAT-AP 3.0.1 (SHACL-conformant, controlled-vocabulary IRIs) for catalog and government interop, and schema.org for open-web discovery. Measure once, emit every dialect.
how we extend it
We do not copy declared metadata into the record. We measure the bytes, add what the standard can carry, and prove the file two ways.
by topology
How the ecosystem encodes each modality in Croissant, and what Meridian emits for it, measured from the bytes.
tabular
in the ecosystem
one cr:Field per column, with a dataType and a source.extract.column recipe
measured by Meridian
the same recordSet, plus the measured range, median, distinct count and null rate as Croissant 1.1 annotations, dataset-level quality on the recordSet (row count, null and duplicate rates) and a Responsible-AI missing-data note from the measured null rate, a record key from the measured identity columns, and a complete categorical vocabulary lifted into an sc:Enumeration the column references
spatial
in the ecosystem
a dataset-level spatialCoverage Place, with no CRS, resolution or band vocabulary
measured by Meridian
GeoCroissant geo/1.0: the measured CRS, pixel resolution with its unit, and band count and names for raster; for vector, a real recordSet of the OGR attribute schema sourced by column. A WGS84 GeoShape box only when the bounds truly are WGS84
tensor
in the ecosystem
native N-dimensional arrays in Croissant 1.1: isArray and arrayShape
measured by Meridian
a 1.1 recordSet with one isArray field per measured array, each carrying its arrayShape and a cr-numeric element dtype (Float32, Int64 and so on) read from the bytes
structure
in the ecosystem
a cr:FileObject for the mmCIF archive, with no molecular vocabulary: a declared blob
measured by Meridian
a measured structure read from the mmCIF bytes: experimental method, resolution, space group, and the counted chains, residues, atoms and element makeup. The 3D coordinates themselves, not a declared archive
modal
in the ecosystem
a cr:FileSet with an sc:ImageObject, AudioObject or VideoObject field sourced by fileProperty content
measured by Meridian
a measured FileObject description: format, width by height, color mode, frames or page count. An honest decline rather than a fabricated single-record field
composite
in the ecosystem
a cr:FileObject archive plus a cr:FileSet of glob includes
measured by Meridian
the archive FileObject with a measured breakdown of its constituent topologies, declining a FileSet we cannot honestly glob yet
Every other Croissant producer copies the publisher's declared metadata into the record. Meridian measures the bytes first, then writes the record: the per-field statistics, the array shapes and dtypes, the coordinate system and pixel resolution, the vector attribute schema in a Meridian Croissant file were all computed from a kilobyte-range read of the actual data, not asserted by whoever uploaded it. Where we did not measure something, we decline to a measured description rather than invent a recordSet. And we validate every file two ways: against the real mlcroissant library on datasets harvested live from Zenodo, and against our own SHACL shapes for the GeoCroissant terms the library does not check.
measure once, speak every dialect
The structure is measured once, from the bytes. From that single source of truth Meridian projects every catalog dialect at the same endpoint, so a dataset is loadable, discoverable, and federatable without re-curation.
schema.org
how you are found
the open-web Dataset vocabulary Google Dataset Search indexes
DCAT-AP 3.0.1
how you federate
the EU catalog-interop profile portals and governments harvest, emitted with controlled-vocabulary IRIs and SHACL-validated
Croissant
how you load
the ML-native contract a training pipeline reads to materialize you
The EU machine-learning profile MLDCAT-AP is tracked as it converges: it builds on the DCAT-AP base Meridian already speaks and points at a Croissant file by conformsTo, so the bridge extends to it without a rewrite.
route to compute
The last mile onto national compute
You hold the loader; here is the path it travels onto national compute, in the order a researcher walks it. Meridian owns one step, measuring the dataset and minting the loader, and routes the data into staging and the run. The identity, the allocation, the credits, the transfer and the node stay with NSF cyberinfrastructure (NAIRR, ACCESS, Globus, CILogon, Open OnDemand): Meridian prepares what runs, never where it runs.
where it sits
The fabric you now live in
Pull back. You now sit inside the national cyberinfrastructure as the measured action layer: repositories upstream, federation alongside, agents and researchers and compute downstream, standards spoken both ways.
connect
You are loadable now. This is the floor.
The same measured engine is live and reachable: an agent adds the hosted MCP endpoint with a bearer token and gets every tool above against the real catalog. You have gone as deep as the bytes go. Everything past this line is the way back up - the question that came down to find you now carries you home as an answer.
claude mcp add --transport http meridian https://meridian.brightquery.ai/mcp --header "Authorization: Bearer $MERIDIAN_TOKEN"The ascent
You rise back up as an answer
This is the return leg, and it is the moat. Byte-level measurement is the bedrock: it is capital, and once it is funded anyone can rebuild it. The moat is the workflow that stands on it - a research method a lab adopts and cannot easily leave. A problem becomes a query plan you approve before it runs, graded evidence you keep or dismiss onto a provenance graph, and an answer where every claim cites what you kept. The descent left you measured and loadable at the floor; the climb turns those bytes back into meaning. You are no longer a file nobody opened. You are a citation someone can trust.
Discover, unified
One query, three families. Byte-measured datasets are the bedrock; scholarly works and patents ride alongside as declared evidence, each graded honestly on its card, each failing on its own so a flaky source never blanks the page. When nothing matches, it says so and names the terms it could not resolve, never padding the page with near-misses.
Literature to data
Open the paper you kept and the byte-measured datasets that cite it rise into view - the reverse of a citation - plus datasets reached by walking the citation graph, each with its traceable path. Paper-first tools stop at the paper; Meridian carries you to the bytes that back it, and one of them is the 4.2 GB file you followed down. The literature, resting on the data. The data was you.
Grounded answers
A question becomes a plan, evidence you keep or dismiss on a provenance graph, and an answer where every sentence cites its evidence and wears its grade. Open a row and its full measured provenance docks beside the answer, inspected in place without leaving it. A firewall keeps and flags any claim that cites what we never provided: never silently dropped, never silently trusted. This grounds an answer in the evidence you kept. It does not synthesize the literature, name its white spaces, or close them, and we do not say it does.
Lifecycle coverage
The scientific data lifecycle
A scientific dataset moves through about ten lifecycle stages, from acquisition to archiving. Meridian is deliberate about which ones it does rather than claiming the whole arc: it owns three - exploration, discovery, analysis - and the other seven belong to the research teams, the repositories, and the partners around it. Naming their owners is the point rather than a concession.
The three we claim, defined
- exploration.
- You name the research problem, not a keyword. The Scout proposes where to look and shows the query it plans to run before it runs anything; you decide what counts as proof. See it in Research.
- discovery.
- What exists anywhere, across measured datasets, scholarly works and patents, each result carrying how we know it: byte-measured, or merely declared by its publisher. See it in Discover, Sources, Quality.
- analysis.
- Reading into the data itself. We open the bytes and measure structure, quality and grain, rather than transcribing a publisher's metadata; and for data that cannot move, we route the analysis to it through OpenMined's Syft instead of moving the data to the analysis. See it in the measured profile on every dataset, and the Syft compute recipe.
The seven we do not
- acquisition
- (research teams).
- Getting data into your hands, licenses included. Meridian never takes custody of bytes.
- transfer
- (research teams).
- Moving files between instruments, institutions and storage. Not ours to move.
- management
- (repositories).
- We observe it. Link health, freshness and drift are reported, never managed.
- curation
- (repositories and their stewards).
- Our completeness and quality scores surface gaps for a steward. Finding a gap is not curating it.
- sharing
- (repositories, and OpenMined · PySyft for data that cannot be shared).
- Repositories remain the system of record. The interesting case is the data that normally cannot be shared at all, and Syft is the mechanism that lets it be.
- synthesis
- (research teams).
- Synthesis means going back to the literature, naming its gaps and white spaces, and closing them. We ground claims in the evidence you kept. That is not the same thing, and we do not claim it.
- archiving
- (repositories).
- We observe it. A dataset that goes dark is reported, not rescued.
What happens in each stage
The ten stages, the activities in each, and who tends them. AI Alliance, the non-profit prime, convenes the work and delegates the data platform to BrightQuery (Meridian); research teams generate the data; OpenMined · PySyft runs the analysis that cannot move the data, and ML Commons carries Croissant, the metadata standard discovery rests on. Repositories keep the rest.
01acquisition
research teams- Scientists run experiments and deploy instruments to collect fresh observations.
- Each reading is recorded in whatever form the instrument or survey produces.
02transfer
research teams- Newly gathered data is moved off instruments and field sites into durable storage.
- Files travel between institutions to where the work will actually happen.
03management
repositories- Every record is organized, described, and kept healthy so it stays usable.
- Routine checks catch errors early and flag anything that has gone stale.
04exploration
BrightQuery- A question opens a workspace: one search across measured datasets, scholarly works and patents, with the literature linked to the data that backs it.
- Evidence is kept or dismissed on a provenance graph, and a grounded, cited answer comes back - every claim tied to what you kept and graded.
05analysis
BrightQuery with OpenMined · PySyft- Methods and models run against data that is too sensitive to leave its home.
- Only the findings travel back; the protected data itself never moves.
06curation
repositories and their stewards- Each record is reviewed for completeness, with gaps and errors written down.
- A reliability score is attached so users know how far to trust it.
07sharing
repositories · OpenMined · PySyft- Datasets are published with clear licenses and rules for reuse.
- Records are indexed so both people and AI tools can actually find them.
08synthesis
research teams- Research teams combine measured datasets and models, in the open, into something new.
- Findings from many studies are drawn together into fresh conclusions, the grade of each input still legible.
09discovery
BrightQuery- Researchers and AI systems search the catalog and weigh quality signals.
- The right dataset is pinpointed, with a clear route to reach it.
10archiving
repositories- Records are preserved well beyond the life of any single project.
- They are re-checked over time so they stay readable and verifiable for years.
Not from zero
Built on deployed systems
A Category-2 solicitation builds upon an existing, deployed system, not a blank page. Meridian is the next generation of two BrightQuery systems already in production for U.S. statistical agencies, carrying the same grounded-evidence discipline from government data to the measured scientific record. These are two of the five foundations it builds off and orchestrates; the rest are the consortium below - Hugging Face, OpenMined, and Croissant, with AI Alliance convening.
BrightQuery Navigator
Deployed for statistical agencies
The assistant that answers questions from government data accurately - grounded in the source record, never a model guessing.
Meridian carries forward the same grounded-evidence discipline, turned from government datasets to the measured scientific record. Meridian is its natural next generation.
BrightQuery DUP
Deployed · Data Usage Platform
Finds where the published literature uses a given dataset - the citation trail from a paper back to the data it rests on.
Meridian carries forward the papers-to-datasets link, the declared reverse index and the tangential citation hop, applied to scientific datasets instead of government ones.
Who builds this
The consortium
Five organizations, and what each actually contributes. Every row carries the state of the integration rather than the intention behind it, because a logo is a claim. The most useful row is the least flattering: Hugging Face datasets are graded declared, since they computed the profile and we read none of the bytes. A project that will not flatter its own partner can be believed about the grades everywhere else.
AI Alliance
PrimePlannedThe non-profit that convenes the consortium and delegates the data platform to BrightQuery. Every partner here was already a member; that was designed, not lucky.
In the code today. Convening and prime award. No code surface.

BrightQuery
Builds MeridianLiveThe prime delegate: the AI Alliance hands the data platform to BrightQuery, which builds Meridian. Two systems already deployed for U.S. statistical agencies are its heritage - Navigator, which answers questions grounded in the source record, and the Data Usage Platform, which traces the citation trail from a paper back to the data it rests on.
In the code today. Meridian IS the build. Its measurement-and-grounding spine and its declared papers-to-datasets index descend directly from Navigator and the DUP, turned from government data to the measured scientific record.
OpenMined
Funded partnerPrototypeSyft: run a method against data too sensitive to move. Only the findings travel; the bytes never do. It is also the only honest answer to sharing, because the interesting data is the data that can never be put in a repository.
In the code today. A SyftBox connector reads the public datasite surface and mints a syft:// compute recipe. It routes; it does not yet measure the published mock, and it is not registered, so no datasite is harvested in production.
MLCommons
Funded partnerLiveCroissant, the metadata standard for machine-learning datasets, already adopted across academia and industry. Discovery rests on it.
In the code today. Every dataset Meridian serves emits Croissant, validated against the shape constraints, alongside full DCAT-AP.
Hugging Face
Deployment platformLiveWhere AI-native datasets and models are published and found, and where what this project builds is meant to land. Not a lifecycle stage, and not a premise of the argument.
In the code today. We index the Hub and bind each dataset from the datasets-server statistics. Those profiles are graded DECLARED, because Hugging Face computed them and we read zero bytes. Their auto-converted Parquet is byte-readable, so this is a grade we can earn rather than a limit we are stuck at.
The analysis stage, in depth
Why Syft, and how we work with it
Analysis is one of the three stages Meridian claims, and it is the one we run with a partner. Meridian's own analysis is the probe: we open the bytes and measure the structure, quality and grain of a dataset rather than transcribing what its publisher declared. OpenMined's Syft extends that to the data we are not allowed to open. It runs methods against data too sensitive to move, so only the findings travel and the bytes never do, which is also the one honest answer to sharing: the interesting case is not the data already in a repository, it is the data that can never be put in one.
why it is vital
The most valuable data is often the data that cannot move: clinical records, federal microdata, proprietary logs. Syft makes it usable without moving it - the code travels to the data, the owner approves and runs it, and only the result returns. It is how the IDSS analysis stage works for data that can never leave home.
how we use it
Syft leaves discovery out of the protocol on purpose (Principle 9: discoverability happens on companion sites like SyftHub). Meridian is that companion: a connector catalogs datasites over their public surface and mints a ready-to-run compute-to-data job as the loader. Today it routes rather than probes: measuring the public mock with the same probe is deferred to Phase 2. Built against the live network.
how we improve it
We give the Syft network the public, measured, AI-native catalog the protocol leaves out, and feed users back into the Syft tools - the minted on-ramp ends inside syft-client. We are aligning with OpenMined on a standard public dataset descriptor so datasites become crawlable and measurable: the missing seam between a private datasite and a discoverable one.
The split is clean and each partner stays in its lane: OpenMined's Syft owns the privacy-preserving compute - the datasite, the mock, the owner approval, the result - and Meridian owns discovery and routing, holding no private data and issuing no approvals. That is the analysis stage of the lifecycle realized: restricted data made findable, assessable, and actionable without a single private byte leaving its owner.
The measurement vocabulary
What a topology is, and the shapes we measure
A topology is the structural shape of a dataset's bytes, not its file format. The format is what a file technically is (CSV, NetCDF, GeoTIFF); the topology is the shape that data takes, and it is what selects the reader. Meridian writes one measurement strategy per shape rather than per format, so a single reader serves every science: CSV, TSV and Parquet are all just tabular. The vocabulary is closed, and honest about its reach: eight shapes are measured by opening the bytes with a purpose-built reader, and two stay declared.
row count, per-column kind, null rate, distinct count, min/max, mean/median/std, percentiles, a 16-bin histogram, monotonicity and duplicate rate (Parquet reads exact stats from the footer)
per-array name, shape and dtype, the array count and a subtype (grid, model weights, vector store); shapes and dtypes, not element values
raster CRS, bounds, width and height, band count and names, pixel resolution with its unit; for vector, geometry type, feature count and the field schema
record count, alphabet (DNA, RNA or protein), min/max/mean length and record ids; variants and alignments read through htslib
node and edge counts, directed and multigraph flags, the node and edge attribute keys; an exact triple count for RDF
experimental method, resolution in angstroms, space group, polymer types, and atom, chain, residue and model counts
image format, dimensions, mode and frame count; PDF page count. Container facts only, never pixel or text content
the member file count and the mix of constituent topologies inside, each member classified by name without opening its bytes
Declared only, no reader yet, so graded declared and never guessed: stream (event and time streams) · unknown (the shape cannot be resolved).
Where the work lands
Why Hugging Face, and how we work with them
Hugging Face is not a lifecycle stage and we do not claim one through them. They are the open platform where AI-native datasets and models are published and found, and they are where what this project builds is meant to land. The hand-off is grade-clean by design: what Meridian opened travels as measured, Croissant-native metadata, and what it only read from the record travels as declared and says so - so a community builds on a provenance it can check rather than a grade it has to guess. That standard is Croissant, and ML Commons carries it.
Why it is infrastructure
Built to be national, not an app
Transdisciplinary
One probe measures a climate NetCDF, a social-science CSV, a geospatial raster and a protein structure through the same ten-topology vocabulary. The mechanism is domain-agnostic by construction, and it shows: the live corpus spans paleoclimate, marine and ice-core geoscience, ML benchmarks and now structural biology.
Operational
Not a demo. A continuously running service: measured structure served over GraphQL, REST and MCP from one set of frozen contracts, live on managed cloud.
Integration-native
It never replaces a repository or hosts the bytes. It sits on top of the ecosystem and federates the cyberinfrastructure around it.
AI-driven
The agent is the customer. Where Meridian opened the file the dataset arrives already measured and already loadable, and where it only read the record it says declared: no bespoke parsing, no guessing, no grade it has to trust on faith.
Proposal figure
One page, for the proposal
The whole system in one page, horizontal or vertical: the research workflow as the moat, the flow you author as data and the trust floor you cannot, the measurement bedrock it all stands on, and the consortium behind it. Flip to doc colors and save it as SVG or PNG.
The return
You are here: cited in someone's answer
You began at the top as 4.2 GB nobody had opened. You went all the way down to your own bytes and came back up measured, loadable, linked to the papers that cite you, and quoted inside a grounded answer with your grade attached. That is the round trip: a question falls to the data and rises as evidence someone can trust. Then the next question arrives, and it starts again.
you are here: measured, loadable, cited
Meridian - the measured-data spine beneath a grounded research workspace.