methodology
How we grade evidence
Every claim on this site about what a compound does carries one of five evidence tiers. This page defines each tier, sets out which sources count and in what order, explains how a compound's A to D grade is derived from its tiers, and says exactly what the review date on a page means.
Why grade at all
Because in this category the strength of the evidence is the story. A claim supported by a completed human trial and a claim supported by a forum thread read the same in a paragraph, and they are not the same thing at all.
Grading turns that difference into something a reader can see at a glance. It also lets us cover claims we could not otherwise cover responsibly. A widely repeated claim with nothing behind it is worth writing about, as long as the page says what is behind it.
The tier attaches to the claim, not to the compound. A compound with a completed trial on wound healing does not get a strong tier on its cognition claim. Each claim is graded from the strongest source that actually tested that claim, in the population the claim is about.
Tiers cover claims about what a compound does, or about what was observed in a body, a cell, or a person. Regulatory, legal and operational facts carry no tier: an agency action or a compounding category is not evidence about an effect, so it cites the source document and its date instead.
Which sources count
In descending order of weight:
- Completed randomized controlled human trials, published in a peer-reviewed venue.
- Other published human studies: single-arm trials, dose-escalation safety studies, observational cohorts, published case series.
- Trial registry entries with posted results, including registrations that show a trial was withdrawn or terminated.
- Published animal studies.
- Published cell, tissue, and computational work.
- Regulatory and governmental records: FDA warning letters and their document numbers, compounding category decisions, the World Anti-Doping Agency Prohibited List, Department of Defense supplement-safety listings.
- Community reports, cited as what people say rather than as evidence of anything.
What does not count
- Vendor marketing copy. Never a source for a factual claim. It can be quoted as an example of what the market claims, attributed as such.
- Other content sites. Never a source. A figure that circulates widely in this category and traces only to affiliate publishers does not get published here, whatever its apparent citation count. If the trail ends at a content site, the number does not ship.
- Conference abstracts without data, for anything beyond noting that work was presented.
- Preprints count, and are labeled as not peer reviewed in the same sentence.
On a compound page or a guide, every claim is cited by number into the reference list on that page, and every reference carries an identifier: a PMID, a DOI, a trial registry number, or a regulatory document number.
How studies are selected
We start from the claim, not from the literature. For each compound we first list every claim area it is actually discussed and marketed for, including the weak ones. Then we search for what tested each one.
- Registry sweep. Every registered trial for the compound, whatever its status, including withdrawn and terminated ones. Status and identifier are recorded even where there is nothing to grade.
- Literature sweep. The primary literature for each claim area, working outward from human studies to animal to cell work.
- Regulatory sweep. The current agency position, with dates.
- Selection. Every human study we find is reported, including negative and null results. For preclinical work, where a claim rests on dozens of similar studies, we cite the studies that define the claim: the earliest, the largest, any that failed to replicate, and any independent replication. We say how many exist and where they came from, because concentration of an entire evidence base in one laboratory is itself a fact about the evidence.
- Negative space. Every page carries an explicit section on what has not been studied and what is not known. It is not an afterthought; it is usually the most useful part of the page.
A study is never excluded for being inconvenient. A trial that failed, was withdrawn, or missed its primary endpoint appears on the page.
The five tiers
Each tier below has a stable anchor that the rest of the site links to.
Fixed vocabulary. Nothing gets added to it without a dated note on this page.
Human RCT
All of the following: human participants, randomized allocation, a control or comparator arm, results published in a peer-reviewed venue or posted to a public registry, and the trial measured the claim being made in the population the claim is about.
Not this tier: registered but not completed; completed but unpublished; randomized without a control arm; a conference abstract without data; a trial of a different formulation or route.
Recorded with the claim: registry identifier, number of participants, arms, duration, primary endpoint, and whether the primary endpoint was met. A missed endpoint is published.
Human open-label
Human participants, with no randomization or no control arm or neither. Includes single-arm trials, dose-escalation safety studies, observational cohorts, and published case series of three or more people. A retrospective chart review qualifies only if it was published with methods.
Not this tier: a case report of one or two people, or a survey of self-reported outcomes. Both are Anecdote.
Recorded with the claim: design, number of participants, duration, what was measured, and the absence of a control arm stated explicitly.
Animal
A peer-reviewed study in a living non-human organism.
Recorded with the claim: species, the model used, route, dose scale relative to body weight, and what was measured. Species and model are mandatory. Improved healing in rats, without naming the injury model, is not a usable claim.
Every animal block on this site closes with a paragraph on what animal work can and cannot show.
In-vitro
Cells, tissue preparations, cell-free systems, or computational modeling. No living organism.
Recorded with the claim: the cell line or preparation, the concentration used, and what was measured. Concentration is mandatory, because concentrations that produce an effect in a dish are frequently far above anything a body reaches.
Anecdote
Community forum reports, self-reported outcomes, podcast and influencer accounts, vendor marketing claims, single case reports, and user surveys.
This tier exists on purpose. Anecdote is what most readers arrived having heard, and a site that refuses to name it sends them somewhere worse. Grading it is what makes it publishable.
Recorded with the claim: what kind of source it is, roughly how widely it circulates, and the standing statement that it is not evidence of effect. We do not reproduce user quotes as testimonials.
Cases decided in advance
| Situation | Tier |
|---|---|
| Registered trial, recruiting, no results | No tier. It is not a claim. It appears in the trials table with its status. |
| Completed trial, results withheld or withdrawn | No tier. The withdrawal is reported explicitly, in the trials table and in the what-is-not-known section. |
| Human trial of a different salt, ester, formulation, or route | Two claims, not one. The claim that trial tested keeps its real design tier. The marketed-route claim gets no tier from that trial, because nothing tested it. The difference is named in the same sentence either way. |
| Human trial of the parent protein a peptide is a fragment of | The same split. Two molecules, two evidence bases. Conflating them is one of the more common errors in this category, and it is how a fragment inherits a trial count it never earned. |
| Systematic review or meta-analysis | The tier of the studies it pooled, never higher. A meta-analysis of rodent studies is Animal. |
| Preprint of a human trial | Human open-label, labeled as not peer reviewed, regardless of design. |
| Approval of a related compound | Not evidence about this compound. No tier. It belongs in the regulatory section. |
| Mechanism inferred from a related peptide | In-vitro or Animal as the source dictates, with the inference named as an inference. |
How the A to D grade is derived
The tier describes one claim. The grade describes a compound's whole evidence base, and it is a summary of the tiers rather than an independent judgment. If the tiers on a page do not support the grade on its card, the grade is wrong.
The four grades are mutually exclusive. Two questions decide them, in order. First: is there a completed human study of this compound? If no, it is D. Second: did any of that human evidence test the use it is actually sold for? If yes, A or B. If no, C.
Evidence about a different molecule never sets the grade. A peptide that is a fragment of a larger protein does not inherit that protein's trials, and neither does a close analog. The related work is reported in the trials table, labeled as being about the other molecule. This is the most common way a compound in this category picks up a trial count it never earned.
| Grade | What it means | Derivation |
|---|---|---|
| A | Completed phase 3 | At least two completed, published Human RCT trials for the marketed use, at phase 3 scale or equivalent, consistent across more than one program. Two trials, not two claim rows and not two citations: claims whose sources overlap are one trial. Every one of them must be cited, must state whether it met its primary endpoint, and they must agree; and every completed randomized trial of this compound for the marketed use must state whether it ran at phase-3 scale, whatever that answer is, so that a trial cannot drop out of the agreement check by saying nothing. A trial that has not read out states no scale, because it is not a claim yet. |
| B | Human data, not settled | At least one completed human study tested the marketed use and read out, and it did not settle the question, which takes two phase-3-scale randomized trials that agree with each other. A B says the evidence exists and stops short; it does not say why it stops short. An unfinished program, trials that disagree and trials too small to conclude from all land here, and the page never asserts which unless the trial records show it. |
| C | Human data, wrong question | Completed human evidence exists for this compound, but none of it tested the marketed use. A different condition, a different route, or a different formulation. |
| D | Animals and anecdote | No completed human trial of this compound, for any use. The claim set about this compound is Animal, In-vitro, and Anecdote only. A D-graded page can still carry a Human RCT chip on work about the parent protein it is a fragment of, because that trial keeps its own design tier and is labeled as being about the other molecule. |
Grade B reaches the card under one of three labels, because “not settled” covers three different situations and a reader seeing the letter alone deserves to know which one. Where every randomized trial of the marketed use missed its primary endpoint, the label reads Human data, endpoint missed. Where separate trials disagree with each other, it reads Human data, results conflict. Otherwise it reads Human data, not settled. The letter records how much evidence there is and how good it is; which way the evidence went travels with the label and the sentence beside it.
Most compounds in this category sit at D. That is a fact about the literature rather than an opinion about the compounds, and it is one of the more useful things a reader can take away.
The Claim Strength Matrix
Every compound page carries one table that puts the whole picture in one place, with four columns:
- Evidence Area. The claim, in the reader's words rather than in clinical vocabulary.
- What Has Been Studied. Species or population, design, number of subjects, and what was measured. The model is named, not just the outcome.
- Evidence Level. Exactly one of the five tiers. No blends, no slashes. Claim areas resting on different evidence get different rows.
- What It Can and Cannot Show. Two clauses: what the cited work supports, then what it does not.
Every claim area a compound is discussed for gets a row, including the ones with the weakest support. Leaving a weak row out would be the same failure as writing only about the positives.
Why mechanism is not a result
A standing passage closes every mechanism and preclinical block on this site, because the same caveat applies every time and a reader deserves it every time rather than once.
what preclinical evidence can and cannot show
Results from animal models are not results in people. An animal model is a deliberate simplification: the injury is created on purpose, the animal is young and healthy, the dose is scaled to body weight in a way that does not translate directly, and the outcome is measured at a fixed point rather than lived with. Cell and tissue work is a further step removed, because concentrations that produce an effect in a dish are often far above anything a body reaches. Mechanism is a reason to run a trial. It is not a substitute for one, and compounds that looked mechanistically convincing have failed in people many times.
What Last Reviewed means
It means a person went through that page on that date and checked four things: that the citations still resolve, that trial statuses are current, that the regulatory section matches the present record, and that every tier still matches its source.
It does not mean the page was rewritten, and it is never moved to make a page look fresh. Whole-corpus re-stamping of review dates is common in this category and it is a manipulation signal with no upside. We re-date pages we actually reviewed, and nothing else.
The visible date always matches the machine-readable dateModified on the same page. If you ever find those two disagreeing, that is a bug and we would like to hear about it.
When a tier changes
- A tier moves up only when a qualifying study is published, cited by number, and added to the page's trials table in the same edit.
- A tier that changes because new evidence landed is a normal update, dated on the page.
- A tier that changes because the previous one was wrong is a correction, published in the log with what it was and what it now is.
- The corrections entry states which of the two it was. The distinction matters and we do not blur it.
The limits of this method
Stated plainly, because a methodology page that claims no weaknesses is not a methodology page.
- Tiering is a judgment. The rules above remove most of the discretion, and the decided cases remove more, but a person still assigns each tier. Disagree with one and write to us.
- Five tiers is coarse. A well-run 400-person trial and a sloppy 20-person trial can both be Human RCT. The surrounding text is where sample size, design quality, and funding source get described, and that text is not optional.
- We are not systematic reviewers. We search thoroughly and we cite what we find, but this is editorial work, not a registered systematic review, and it does not carry that guarantee of completeness.
- Absence of evidence is not evidence of absence. A compound at grade D has not been shown to do nothing. It has not been tested. Those are different claims and the site keeps them apart.
- No reviewer with clinical credentials, yet. The organization is the reviewer of record. We would rather say that than invent one, and the editorial policy explains what changes when that is no longer true.