NGARi · Sovereign environmental research
The land keeps a record. So does the community.
Environmental research now runs on two kinds of record that behave nothing alike. A soil core is a measurement: repeatable, citable, and expected to be shared. An oral history, a maroon route, a family's account of a cane-burning season is knowledge held by someone — and the moment it is flattened into a column of numbers it stops being what it was. Most research data systems cannot hold both. They either treat everything as data, or they keep the qualitative material in a folder and the promise to protect it in a sentence on a web page. We build the third thing: one custody chain that carries both, with the consent decision enforced per record.
On-premises · deterministic · no trained model and no weights · no field deployment claimed
Two records, one programme
An environmental research institution studying a landscape shaped by extraction holds material in three incompatible states at once. Getting them into one chain of custody is the infrastructure problem — and it is a governance problem before it is a technical one.
Measurement
Soil chemistry, stable isotopes, microbial and plant diversity, elevation, climate. Quantitative, repeatable, and lawfully publishable to the extent the collection terms allow.
Held knowledge
Interviews, oral histories, archival records, place knowledge. Authored, attributable to a contributor, and governed by a decision that only that contributor or community can make.
Custody
Who collected it, when, under which protocol version, who handled it after, and whether the record can still be trusted a decade later when a claim is challenged.
What the engine does
Six things, all deterministic and all readable by a research team without trusting a model. The hard part is not the statistics — it is the refusal.
An enforced governance gate
Every record carries a sovereignty basis and a consent state. Records with a protected basis, or with consent that is not recorded, never enter an analysis — and are refused by the release policy for every use, including internal ones. Refusal is a decision with a code and a reason, never a silent drop. A record that is quietly skipped is a record somebody will report on anyway.
Qualitative knowledge cannot become a column
Knowledge records carry no numeric proxy, and the engine refuses any that does. Converting an interview into a variable is a specific, common, and irreversible reduction; it is refused before any other test runs.
Association ranking that refuses two things
Ranks soil and landscape variables against ecological response markers. A variable that never varied returns an undefined coefficient, not zero — "no information" and "no relationship" are different findings. Missing values are reported with their true n and are never imputed.
Drift with its confounder named
Compares a baseline window to a recent one and reports what moved. When the protocol version changed across that boundary, the row is labelled as confounded rather than published as an ecological trend. An unreported method change is how a laboratory artefact becomes a finding.
Custody gap report
Walks the declared chain from collection to archive and names what is missing: no collector, no collection date, no method version, a broken chain, an unassigned step, a withdrawn consent still present. Shaped on the ALCOA+ record-integrity vocabulary and on the FAIR and CARE data principles together — because a well-formatted record is not a well-governed one.
A receipt rendered from the ledger
The evidence ledger is hash-chained and append-only. The receipt is rendered from it, never assembled alongside it, so the two cannot disagree. Editing an entry breaks verification; a ledger rewritten end-to-end to verify against itself is caught by its fingerprint.
The honest line
There is no trained model and no weights in this engine, and it has never seen any partner's data. It is a deterministic reference implementation over a synthetic fixture. Every capability a research partnership would need is carried in the engine's own status contract as achieved: false, measured: false, and a guard test reads the source so no future branch can quietly flip one true. What exists is the mechanism and the refusals; what does not exist is a deployment.
What it found in a synthetic fixture
The reference fixture is deliberately ordinary and deliberately imperfect. It is synthetic data — fourteen samples across five sites, written to encode situations a real programme meets — and the engine's output is reproducible from it. Nothing below is a finding about a real place.
| Stage | Result on the synthetic fixture |
|---|---|
| Analysis gate | 11 samples admitted, 3 refused — one protected community-knowledge record, one restricted record, one whose consent was withdrawn |
| Association ranking | Top pair soil_org_c_pct ↔ microbial_shannon_h, r = 0.998 (n = 9). The variable that never varied returned undefined, not 0.0 |
| Drift | clay_pct moved z = +2.46 across the window — reported and tagged confounded, because the method version changed at the same boundary |
| Custody | 7 findings, 4 of them high: a broken chain, an unassigned step, a missing timestamp, and two protected records present in the analysis fixture |
| Governance | 35 uses permitted · 15 refused · 4 escalated to a human. Commercial and third-party transfer refused as a standing posture, not as 60 record-level findings |
| Ledger | 8 entries, chain valid, receipt rendered from it |
The r = 0.998 is a statistic over nine synthetic points. It is not evidence about any soil, anywhere. It appears here because a capability page that only showed the machinery and never the output would be less honest, not more.
What ships, and what is an engagement
| Layer | Status |
|---|---|
| Appliance, air-gap mode, hash-chained audit, signed attestations | Shipping |
| Governance gate, ranking, drift, custody report, ledger and receipt | Working reference implementation, synthetic fixture |
| Read-only connectors for a field notebook, a laboratory system and an archive | Reference adapters; no institution has ever been contacted |
| Offline field capture with lossless reconciliation | Described, not solved — no sync layer exists |
| Open-weight models on institution-controlled hardware | Not started — no model, no weights |
| Public methods record readable without a data-science background | Not started |
| Deployment on a research site | Engagement |
What "sovereign" means here, precisely
It is a word that gets used loosely, so this is the narrow version. Three claims, each of which can be checked:
The records sit on hardware the institution controls
Field and laboratory records are held on an appliance on the institution's own network, not in a vendor's account. There is no per-inference bill and no third-party endpoint holding the record of a landscape.
The decision is the community's, and it is enforced
The engine does not decide what may be shared. It records the decision that was made, applies it per record, and refuses what the decision forbids — including refusing to analyse material that was never offered for analysis.
The evidence can be re-derived
A published figure traces back through a custody chain to a sample and a protocol version, and the chain can be recomputed. That is the difference between a result and an assertion.
What we are not claiming
- No trained model, no weights, no foundation model, and no measured analytical accuracy of any kind.
- No field deployment, no live field notebook, laboratory system or archive, and no partner data of any kind — we hold none.
- No offline sync layer. The field case is described in the code and not implemented; saying otherwise would be the first thing to break in a rainforest.
- No data-governance determination for any community. The engine enforces the decision a community records; it does not make it, and it refuses to guess.
- No land-management recommendation, no regulatory certification, and not legal advice. The custody report is a gap report against a named checklist.
- No claim that any figure in a synthetic fixture says anything about a real ecosystem.
Tell us what your programme has to hold.
If your research depends on material that cannot be published, the governance question comes before the model question. That is the right place to start, and we are happy to start there.