← Insights

Negative Results Are Assets: Turning Failed Experiments Into Licensed Data

Null results are buried on local servers while other labs repeat them — yet verified negative data is exactly what predictive drug-discovery models need to learn from.

16 August 2026 10 min read DimenChain
Figure 1The highlighted cluster is the null result: the densest data, and the least likely to be published.
€250K A rigorous null trial
€750K Spent repeating it
€0 Recovered today

Every laboratory holds knowledge that the rest of its field does not have, and a considerable share of that knowledge is negative. The compound that did not bind. The endpoint that was not met. The catalyst that performed exactly as well as the control. This information is expensive to produce, it is frequently as rigorous as the positive findings that surround it, and in the great majority of cases it never leaves the group that generated it.

The reason is structural rather than cultural. Academic journals rarely publish papers on failed compounds or null endpoints. A null finding does not advance a career, does not attract citations, and does not fit the narrative shape that peer review has been trained to reward. So the data is written up internally, archived on a departmental server, and forgotten. Meanwhile, in another country, a second group forms the same hypothesis, writes the same grant application, and spends the same money finding out the same thing.

The cost of a filing cabinet

Consider a clinical trial of the kind that runs constantly across European academic medical centres. Roughly €250,000 of funding. Well designed, properly powered, ethically approved, competently executed. The compound misses its primary therapeutic endpoint. The trial is, by the conventional definition, a failure.

What that trial actually produced is a substantial and internally consistent body of evidence: baseline toxicity reports, longitudinal biomarker panels, participant genetic profiles, a documented dose-response relationship for a molecule whose behaviour in humans was previously unknown. None of it is wrong. All of it is real measurement. And because the headline result was negative, essentially all of it is filed on local servers and generates no return of any kind — not scientific, not financial, not reputational.

The second-order cost is worse than the first. Other laboratories, in other countries, with no way of knowing what has already been established, run trials on the same failed compound. Public money is spent to reproduce a known negative. Participants are enrolled and exposed to an intervention that a colleague elsewhere has already shown does not meet its endpoint. That is not merely inefficient; it is an ethical problem, and it follows directly from the fact that the first result was never discoverable.

Estimates of the total scale — duplicated experiments, unpublished failures, and the absence of infrastructure through which groups could have found one another — put annual research waste in the order of €157 billion. Whatever the precise figure, the mechanism generating it is not in dispute. Negative results are produced continuously and captured almost never.

Machine learning has a specific appetite for failure

The argument that negative data has commercial value is often made in general terms, and is unconvincing that way. The concrete case is narrower, stronger, and comes from computational drug discovery.

A predictive model that ranks candidate molecules against a target is, mathematically, drawing a decision boundary. To locate that boundary it needs examples on both sides of it. The published literature supplies one side generously and the other side hardly at all, because the literature is a record of what worked. A model trained predominantly on successes learns what successful molecules look like; it does not learn where the effect stops.

The standard workaround is to generate computational decoys — molecules assumed to be inactive because nobody has reported them as active. This is a substitute for evidence rather than evidence. Decoys encode the assumptions of whoever selected them, and models trained against them tend to learn the selection procedure as much as the underlying chemistry. A verified experimental negative is a different class of object entirely: this molecule, against this target, in this assay, at these concentrations, under these controls, produced no measurable effect.

That is why an artificial intelligence firm building predictive drug-targeting models will license a verified negative dataset. It is not buying a disappointment. It is buying the half of the training distribution that the publication system does not supply, and the accuracy of its predictions depends on having it.

The counterfactual to licensing a negative dataset is not open publication. It is a filing cabinet.

The same logic applies outside machine learning, in a blunter form. A pre-clinical programme costing in the region of €500,000, abandoned when a corporate sponsor changed direction, retains its full technical content on the day it is shelved. Sold as a complete intellectual property package for something like €50,000, it allows a startup to skip roughly two years of preclinical work. The seller recovers value from a sunk cost. The buyer avoids repeating work that has already been done properly once. Neither party is doing anything exotic; they are doing what every other industry does with assets it has stopped using.

ElementPublishedHeld back
Study metadata and endpointsPublic—
Proof of ownershipHash on-chain—
Raw trial datasets—Encrypted, off-chain
Participant-level data—Never leaves the institution
Licensing termsIn the contract—
Table 1 — What is published and what never leaves the institution.

The mechanics of listing a negative result

The obstacle has never been that nobody wants this data. It is that a laboratory has had no mechanism to make a dataset discoverable without disclosing it, prove it existed before a given date, and set enforceable terms without a six-month negotiation between two technology transfer offices. That is the gap DimenChain addresses.

What is published and what is not

The separation is strict. Raw data never goes on-chain. What is recorded on Arbitrum is encrypted metadata, cryptographic hashes that prove the existence and integrity of the underlying files, and the events emitted by the licensing contracts themselves. The dataset stays off-chain, encrypted, and under the control of the group that produced it.

The listing process is therefore short:

  • Hash the protocol and the resulting dataset and timestamp both on-chain, establishing an immutable record of what was done and when.
  • Publish descriptive metadata: target, compound class, assay type, sample size, endpoint definition, control structure, instrumentation, date range. Enough for a counterparty to judge relevance and rigour; not enough to constitute the data itself.
  • Keep the raw dataset encrypted off-chain, released only on execution of a licence.
  • Set the commercial terms in the smart contract — price, field of use, duration, whether the licence is exclusive, whether academic access differs from commercial access.
  • Record contribution with Soulbound Tokens, so that authorship of the underlying work stays attached to the researchers rather than to whichever institution employed them at the time.

The timestamped protocol hash deserves emphasis. A negative result registered before the experiment ran is materially different from one assembled afterwards, and the chain makes that distinction verifiable to a buyer who has no other reason to trust the seller.

The objections, taken seriously

Does monetising negative data undermine open science?

Part of this objection is correct and should be conceded plainly. If a null result carries a price, some groups will withhold or delay disclosure that they would otherwise have made freely. That is a real cost and it should not be argued away.

But it has to be weighed against the actual status quo rather than an idealised one. The overwhelming majority of negative results are not openly published today. They are unavailable at any price, to anyone, forever, at zero benefit to science. Against that baseline, a system in which the existence of a result is public and only the raw data is gated is a substantial improvement — because the duplication problem is solved by discoverability, not by access. A group that consults the metadata layer and finds that the compound has already failed a properly powered trial does not need to buy the file to avoid wasting two years. It only needs to know.

Licensing terms are set by the originating group, not by the platform, and nothing prevents terms that are free for academic reuse and priced for commercial reuse. Institutions with open-science mandates can honour them; institutions that need to recover cost can do that instead.

Consent, special category data and GDPR

This is where care matters most, and where the honest answer is a constraint rather than a capability. Clinical datasets containing biomarker panels and participant genetic profiles are special category personal data. What may lawfully be licensed is determined by the consent participants actually gave and the legal basis under which the data was collected — not by what a smart contract says. If the original consent did not extend to secondary commercial use, no contractual mechanism creates that permission retrospectively. In many trials it will not, and the correct outcome is that the dataset is not listed, or is listed only in a form that falls outside the scope of personal data.

It is also worth stating clearly that anonymisation and aggregation are not automatic safe harbours. Genomic data is intrinsically resistant to anonymisation, and aggregated metadata drawn from a small cohort can be re-identifiable in combination with other sources. Pseudonymised data remains personal data under the GDPR. Any assertion that a clinical dataset has been rendered non-personal is a factual claim requiring assessment, and it belongs with the data controller and their data protection officer, who hold the responsibility and the necessary context.

The architecture is built with these constraints in view. Keeping raw data off-chain is not only a confidentiality measure; it is what makes rights of erasure and rectification exercisable at all, since those rights cannot be honoured against an immutable ledger. What the chain holds is a hash: proof that a file existed at a given time, from which the file cannot be reconstructed. Where clinical data cannot be licensed, derived and non-personal outputs frequently can: assay-level results, compound behaviour, toxicity signals detached from individual records. That distinction is the work, and it is legal work, not a technical shortcut.

A null result from a badly powered study is not an asset

Entirely correct, and it is the objection most likely to determine whether a market in negative data is useful or actively harmful. There is a categorical difference between “this compound does not work” and “we were unable to detect an effect”. The second is a statement about the study, not about the world. Feeding underpowered nulls into a training set does not improve a model; it teaches it noise.

Which is why the metadata layer carries what a methodologist would ask for first — sample size, power, control structure, assay validation, the pre-registered protocol and its timestamp — and why that information is published rather than summarised. Rigour is not certified by the platform; it is disclosed, and priced by the counterparty. A buyer training a predictive model has every incentive to distinguish a well-powered null from an inconclusive one, and to pay accordingly. Verifiable provenance is what makes that discrimination possible.

What changes

None of this requires a change in scientific values. It requires a change in what happens to a dataset on the day a project ends. At present the default is archival and silence. The alternative is a record that the work was done, held with cryptographic proof of when and by whom, described in enough detail that others can avoid repeating it, and available under terms its authors set themselves.

DimenChain operates from the Netherlands, is built to be GDPR-aligned by design, and its underlying method is filed under WIPO PCT/IB2024/058791. The platform is in production. Laboratories hash and timestamp protocols, datasets and milestones on Arbitrum, govern collaboration through smart contracts, and licence unpublished and failed research to counterparties who need it.

A failed experiment is still a measurement of reality. The only question is whether anyone else ever finds out.

See this on your own research

A short walkthrough with the team, using your institution’s actual workflow.

Log in →