Insights  /  Original Research

What 80 NIH-funded biotech websites do not say

I assessed 97 biotech companies holding active NIH SBIR or STTR awards against the five things an investor or business development reader is trying to establish. Eighty of the sites could be assessed. Not one scored low risk overall.

That is not a rhetorical opening. Across five criteria and 400 individual judgements, the assessment returned 49 companies at high risk, 31 at medium, and zero at low.

What is more interesting is where the failures cluster. On two of the five criteria, these sites do well. On the other three, across 240 judgements, there was exactly one low risk result.

What was measured

Five things, chosen because they are what a generalist investor, a licensing manager, or a business development analyst is trying to establish when they open a company website for the first time.

  1. Development stage. Can a reader tell where the programme actually is? Pre-IND, Phase I, TRL 4, first in human planned for a named period.
  2. IP position. Can a reader confirm the technology is protected without calling you? An issued patent number, a disclosed pending application, a named exclusive licence, a trial registry identifier.
  3. Business development and licensing. Can a partner tell what you want from them and how to start?
  4. Commercial case. Is the value proposition legible, or is it buried in dense academic language?
  5. Indication and modality. Can a reader place you against comparable companies? Is the indication named, the modality stated, and the difference from similar companies explained?

Each criterion returns low, medium or high risk, and every finding is tied to a line quoted from the company's own pages. Nothing is inferred from a company's reputation, its funding history, or its science. Only what the site says.

The split

Here is the distribution across the 80 sites that produced an assessment.

Criterion results across 80 assessed sites
Criterion Low risk Medium High risk
Commercial case legible223424
Indication and modality named95120
Development stage stated02654
IP position checkable03545
BD pathway visible12554

The top two criteria behave like a normal distribution. Some sites do this well, most do it adequately, some do it badly. That is what you would expect from any sample of anything.

The bottom three do not behave like a distribution at all. They behave like an absence.

Companies are, on the whole, capable of saying what they do. Twenty-two of eighty state their commercial case clearly enough that a nonspecialist reader gets it on the first pass. Fifty-one name their indication and modality specifically enough to be placed against comparable work.

The same companies almost never say what stage the programme is at, what they own, or how a partner would begin a conversation.

Why this is an absence and not a harsh grade

A reasonable objection at this point is that the assessment is simply strict. If the bar is high enough, everything fails.

The data answers this directly, because every judgement records whether the assessment found supporting evidence on the page or found nothing and cited the nearest related text instead.

On IP position, 84 percent of findings rest on nothing. On business development, 84 percent. On development stage, 72 percent.

Compare that with the two criteria where companies do well: 32 percent and 25 percent.

So the finding is not that these companies disclose their IP position weakly. It is that on 84 percent of the sites assessed, there was no IP disclosure anywhere on the pages a reader reaches, and the assessment had to cite the closest thing to it.

The most common form of this is a phrase every reader of biotech websites will recognise. Proprietary platform. Patented technology. Patent-pending process. No number attached to any of them.

The patent usually exists. It is the single most fixable finding in the entire dataset, because the work has already been done and paid for, and it simply never made it onto the website.

What the award says and the website does not

Every company in this sample has a public NIH award with a public abstract. That abstract states, in the company's own words, what they told a federal funding agency they were building.

So there is a second comparison available that has nothing to do with judgement: does the website state the thing the award describes?

Of the 79 companies where this could be checked, 19 were aligned. Sixty were not. Seventy-six percent of these companies told NIH something that appears nowhere on their own website.

Breaking the absent items down by what kind of thing they are:

Absent items by category
What is missing Share of cases where it was checked
Population served66 percent
Molecular target62 percent
Indication50 percent
Modality or technique44 percent
Named asset or product20 percent

Read that ordering carefully, because it is the opposite of what you would guess.

Companies are best at naming their asset and their technique. They are worst at naming who the thing is for and what it acts on.

A company will tell you it has built a microfluidic assay. It will not tell you the assay is for a specific patient population in a specific care setting, even though it told NIH exactly that eighteen months ago in order to get the money.

Concrete examples from the sample, anonymised:

  • An award for at-home hormone testing in perimenopause. The word perimenopause appears nowhere on the site.
  • An award naming a specific molecular target as the mechanism of action. The site names the drug candidate on eight pages and the target on none.
  • An award for a diagnostic in a named infectious disease. The site describes the platform in general terms and never names the disease.

None of these companies is hiding anything. In every case the science is real and the award is public. The information simply did not survive the trip from the grant application to the website.

The 17 percent that could not be assessed

Seventeen and a half percent of the sample produced no assessment at all.

Eight sites could not be fetched. Four of those return an HTTP 403 to any automated request, which is a hosting configuration choice rather than a statement about the company. Eight produced content the assessment refused to score. One was rejected as too thin to assess.

That last category deserves its own note. The threshold for refusing to assess a site is 150 words of extracted text across the entire site. One company in this sample did not clear it.

Separately, of the 80 sites that were assessed, half were assessed on two pages or fewer, because that is how many pages of substantive content the assessment could find. Sixteen companies have a single page.

Method

Population. 206 companies sourced from NIH RePORTER across molecular diagnostics, PCR, infectious disease, point-of-care testing, antimicrobial resistance, protein engineering and sequencing. Of those, 141 have a company website on record with a resolving domain and a public award abstract.

Sample. 97 drawn by seeded random sample from those 141. The seed is stated so the draw is reproducible. The sample was not ordered by any quality signal, because the claim is about SBIR-stage biotech sites in general rather than about companies that looked promising.

Exclusions. Three companies were removed from the frame before analysis because their website is a hosting platform page rather than a company site, one on LinkedIn, one on Google Sites, one on Strikingly. This exclusion is defined on a property known before any assessment ran, not on how those companies scored. Running the analysis with them included changes no figure by more than one company.

Assessment. Each site is fetched, up to three pages of substantive content are read, and the five criteria are scored against the extracted text. Every finding must cite a passage that exists verbatim on the page. Quotes are lifted programmatically rather than generated, so a finding cannot be supported by text the site does not contain.

Award comparison. The most recent award abstract for each company is compared against the site text. Up to four subjects are identified from the abstract, and each is checked against the site by exact match rather than by judgement.

Date. All sites were assessed on 24 August 2026.

What this does not show

The sample is drawn from companies I sourced for my own purposes, in specific therapeutic and platform areas. It is not a random draw from all SBIR awardees, and the findings should not be generalised beyond life-sciences companies of roughly this stage.

Half the assessed sites were read on two pages or fewer. A finding of high risk on such a site means the information is not on the pages a first-time reader reaches. It does not prove the information is absent from every page of the site. That said, an investor opening a company website for the first time reads roughly the same pages.

Seventeen and a half percent of sites could not be assessed, and those failures are not random. A site that blocks automated requests may be a more professionally built site, which would mean the sample underrepresents the better end.

Corpus sizes across the assessed sites span two orders of magnitude, from 13 words to 2,133. A site with almost no text and a site with substantial content each count as one observation.

I also tested whether companies approaching the end of their award funding present themselves any better, on the theory that they would be actively fundraising. They do not appear to present themselves worse, and the sample is too small to say more than that.

What I take from it

The gap is not a writing problem. These are companies with real science, public funding, and in most cases granted or pending IP. The information an investor needs exists. It sits in the grant application, the patent filing and the founder's head, and it does not reach the website.

There is a reason for that, and it is not carelessness. An SBIR-stage biotech website is usually written by the people who did the science, for readers who already understand it. That is the correct instinct and the wrong audience. The reader you need is a generalist who has forty other companies to get through this week, and who will leave with whatever your homepage told them.

If you want to know where your own site sits against these five criteria, I run the assessment and send the write-up. There is no charge. That is the BioSite Audit.

The anonymised dataset is available as a CSV: 97 companies, per-criterion results, award comparison outcomes, and page coverage. Companies are identified by key only. Award numbers, end dates and the absent terms themselves are withheld, because each of those identifies a company against the public NIH record.

Wondering where your site sits?

Run the same five criteria on your own site

Send the URL and your award abstract. I run the assessment and send the write-up, at no charge.