Harell Data Raises $15M for Scientific AI Data

Harell Data, a Bellevue, Washington startup founded by Adaptive Biotechnologies scientific founder Harlan Robins, raised $15 million on August 17 to build a cloud service connecting AI model builders with proprietary scientific training data. Timmerman Report and GeekWire report that Fuse, Cercano Management, and other investors participated. The company is initially working with datasets from A-Alpha Bio and Adaptive Biotechnologies.
Harell Data, a Bellevue, Washington startup founded by Adaptive Biotechnologies scientific founder Harlan Robins, raised $15 million on August 17 to connect AI model builders with proprietary scientific datasets, according to Timmerman Report and GeekWire. Fuse, Cercano Management, and other investors participated in the round.
Robins told GeekWire that Harell's goal is to connect AI modelers with proprietary training sets for difficult scientific problems. He argued that organizations generating such data often lack a commercially viable and secure way to share it with model developers.
Initial data and compute partners
According to Timmerman Report, Harell has arranged secure cloud computing through CoreWeave Cloud using Nvidia GPU technology. The service is initially being populated with two proprietary datasets:
- •A-Alpha Bio data on protein-protein binding affinities.
- •Adaptive Biotechnologies data on T cell receptor binding.
Robins told Timmerman Report that data contributors would receive a share of revenue Harell earns from AI model builders. He also said model builders could obtain curated, drug-discovery-relevant data through the service.
The data-access problem
Robins cited AlphaFold's use of publicly available experimentally derived protein structures as an example of how long-running scientific data collection can underpin major AI advances. He told GeekWire that many biological and medical datasets require years of work and millions of dollars to generate, while their owners may be unable to monetize them effectively.
Harell's proposed model addresses a central constraint in scientific machine learning: data quality, provenance, licensing, and controlled access can be as consequential as model architecture. Companies attempting comparable data-sharing arrangements commonly need mechanisms for governance, access control, compensation, and reproducibility, particularly when training data has commercial sensitivity.
The reported launch is notable for drug-discovery teams because protein interaction and immune-receptor binding datasets are difficult to generate experimentally and can be highly relevant to model training. Public reporting has not detailed pricing, data-access terms, model-training restrictions, or technical controls for governing downstream use of contributed datasets.
Key Points
- 1Harell Data raised $15 million to connect proprietary scientific datasets with AI model builders through a revenue-sharing cloud service.
- 2Initial datasets cover protein-protein affinities and T cell receptor binding, two experimentally demanding biological data types relevant to drug discovery.
- 3Comparable scientific-data marketplaces depend on governance, licensing, provenance, and secure compute alongside access to high-quality training examples.
Scoring Rationale
The funding round targets a meaningful bottleneck for life-science ML: access to curated proprietary datasets with clear commercial terms. Its immediate scope is narrow, but the approach is relevant to teams training scientific models where public data is insufficient or poorly matched to a task.
Sources
Public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

