73,352 parameters
Small enough to run on commodity hardware and deploy inside an existing research environment without new infrastructure or procurement.
Origin Neural built a complete computational discovery catalogue and the engine that produced it. The work is finished, enriched with ADMet and drug-likeness profiling, and signed and anchored to a public blockchain — so you can confirm independently, without our involvement, that this data predates our first conversation and has not been altered since.
The engine is compact enough to run on infrastructure you already own. No supercomputer purchase required.
These are final figures. The catalogue was sealed on 22 February 2026 and has not changed since.
Origin Neural is bootstrapped. We funded this work ourselves and ran the pipeline until we had a catalogue worth licensing, then stopped generating new predictions because generating at volume is expensive and we chose not to burn capital producing inventory before securing a partner. Model development never stopped — training runs continue and are signed and anchored daily, which anyone can verify on-chain. The engine is intact and restarts on demand, against your targets, on your schedule. Restarting it is one of the things a partnership pays for.
The industry trend is toward very large in-house compute. Our approach went the other way: a deliberately small network, trained for geometric stability, that produced this entire catalogue on a bootstrapped budget.
Small enough to run on commodity hardware and deploy inside an existing research environment without new infrastructure or procurement.
Measured on data held out from training. Self-reported — we recommend independent benchmarking against your internal reference sets before any commitment.
A near-zero measure of weight-manifold stability across five training cycles, with isometry loss of 0.386 relating to chirality and rotational robustness.
A mathematically stable model is not automatically a clinically accurate one, and a computational prediction is not a drug. These metrics describe how the model behaved, not whether a compound will work in an organism. We expect and welcome independent validation.
Five distinct offerings. Most partners start with one and expand.
The existing predictions, in full or scoped to a target family, disease area, or scaffold class. Includes compound dossiers, SMILES, ADMet profiles, and scaffold analysis with bulk export.
Deploy the prediction engine inside your own environment and run it against your proprietary targets. Your data never leaves your infrastructure.
Provide a biological target; we restart the pipeline and generate a new, dedicated prediction set to your specification, anchored and delivered.
Anchor your own computational results to establish provable priority dates and tamper-evident audit trails — for IP disputes, regulatory records, or multi-party collaborations.
Joint work on a disease area, target class, or compound family, with terms structured around shared discovery rather than a flat data purchase.
We can prepare a sample set against one target of your choosing so your team can assess quality directly, before commercial terms are discussed.
A screening funnel only has value if you can see where the filters land. These are the real distributions, published so your team can judge fit before a conversation rather than after one.
| Measure | Value | Note |
|---|---|---|
| Peptide-like | 21,151,570 | The majority of the catalogue |
| Drug-like | 12,150,319 | Conventional small-molecule space |
| Fragment | 24,027 | Fragment-based starting points |
| Lipinski compliant | 31.8% | Of enriched compounds |
| PAINS clean | 39.5% | Free of common assay-interference motifs |
| Veber pass | 3.6% | Stricter oral-bioavailability criteria |
| Lead-like | 0.3% | Strictest filter; approximately 100,000 compounds |
| Avg. synthetic accessibility | 4.34 | Scale of 1 (easy) to 10 (hard) |
Roughly two thirds of the catalogue is peptide-like rather than conventional small-molecule, and the strictest lead-like filter passes 0.3%. If your program is exclusively small-molecule oral, the relevant subset is the 12.1 million drug-like compounds, not the headline 33 million. We would rather you know that now than discover it in diligence.
| Classification | Compounds | Meaning |
|---|---|---|
| Green | 3,108,279 | Few structural concerns identified |
| Yellow | 3,595,212 | Caution flags identified |
| Red | 10,491,891 | Significant structural or toxicity concerns |
Every batch was hashed into a Merkle tree, wrapped in an ECDSA-signed payload, and written to the Bitcoin SV blockchain as it was produced. Those records are public and permanent. Anyone can read them without our cooperation.
Block timestamps are established by network consensus. A record in block 937,482 cannot be backdated afterwards by us or anyone else.
Changing any single prediction changes its batch's Merkle root, which then no longer matches the immutable on-chain value. Silent revision is impossible.
Each payload carries an ECDSA signature, so the record was not merely posted to the chain but authored by a specific key holder — and that is checkable.
We publish the raw anchors, the exact payload structure, and a dependency-free script that recovers the signing key and checks it against the address the payload claims. Read the script before you run it — it is short enough to audit.
Open the Verification PageIf you already run large in-house compute, you may not need the catalogue — but you may still want the engine for its efficiency, or the verification layer, which is not something compute alone provides. If you do not run that infrastructure, this catalogue represents work already paid for and independently timestamped, available immediately rather than after a procurement and build cycle.
Cost. We are bootstrapped and self-funded this work. Rather than continue spending on compute to grow inventory we had not yet monetised, we sealed the catalogue and turned to partnerships. The engine is intact and restarts on demand.
Not yet by a third party. Our reported holdout accuracy is 85.40%. We consider independent benchmarking against a partner's internal reference sets a reasonable precondition for any significant agreement, and we will support and fund it.
It is a model-based estimate of interaction strength with a biological target, intended for ranking and prioritisation. It is not evidence of medical effectiveness and should not be read as such.
Yes. We will prepare a scoped sample against a target you nominate so your computational chemists can assess it directly. We would rather be evaluated on output than on claims.
Negotiable and defined per agreement. For commissioned campaigns and engine deployments run against your proprietary targets, our expectation is that the outputs are yours.
Tell us a target and what your program needs. We will come back with a sample set and an honest assessment of whether this catalogue fits — including if it does not.