Novellia
Real world data company built on records patients choose to contribute, indexed for the life sciences data platform rather than the free consumer application, which is the collection mechanism rather than the product sold. Patients use a free app and web platform to aggregate their records from more than 50,000 US healthcare providers across Epic, Oracle Health, athenahealth, Quest, and the VA, then consent to contribute de identified data for research.
Novellia's AI stitches those fragmented sources into structured longitudinal patient journeys spanning a reported 15 to 20 years, which the company positions against conventional real world data assembled by brokers from claims and partial hospital records. Buyers are biopharma teams in HEOR, market access, clinical operations, and medical affairs; the company reports customers among a majority of the top 10 to 15 global pharmaceutical companies. Raised an $18 million Series A led by Spark Capital in June 2026, bringing total funding to $28 million. Founded by Shashi Shankar, previously at Genentech working on real world data.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
The AI does the load bearing work of the data product: stitching records from more than 50,000 provider sources into structured longitudinal journeys and cleaning them to surface signals that snapshot datasets miss. Held back from A because the differentiated asset is the patient consented collection channel itself. A competitor with the same models but purchased broker data would have a materially worse product, which means the moat is provenance as much as inference.
An oversight mechanism that is structurally novel in this index, and incomplete in one important direction.
The system syncs and stitches automatically. What makes this record different is who reviews the result. The company states that patients can see their consolidated history, correct outdated information and add detail themselves. The data subject is therefore the reviewer, and the data subject is the one person in the world who can tell whether a record has been merged onto the wrong patient, whether a diagnosis is stale, or whether an entry is simply wrong. No other real world data model in this index puts the person the data is about in that position, because in broker assembled data the person never sees it.
That is a genuine control and it is worth stating clearly, because it addresses the failure mode that most concerns a research buyer.
What is incomplete is the direction of travel. Nothing published states whether a correction made after data has been contributed to research propagates to the contributed dataset, what happens to analyses already run on the uncorrected version, or what checks a merge before contribution for patients who never review their record. Review by the data subject only works for records the data subject actually reads, and a twenty year history assembled in thirty seconds is a great deal of material to read.
Ask whether corrections propagate downstream, what proportion of contributed records have been reviewed by the patient, and what automated checks apply to merges that nobody reviews.
The capability is clearly described and the technology is not.
What is public: an AI powered platform that connects to more than fifty thousand provider sources, pulls records automatically once portals are linked, accepts paper documents through the camera, and assembles the result into a structured longitudinal journey reported at up to twenty years and completed in around thirty seconds. That is a clear functional description and the speed claim is a meaningful operational fact.
What is absent is everything beneath it. No model or method is named, no architecture described, no account of how record linkage and entity resolution are performed, no handling of conflicting entries between sources, no extraction accuracy for documents captured by camera, no versioning, and no evaluation of any component.
For this product the linkage layer is the whole technical proposition, because the differentiator is coherence across sources rather than volume. A buyer purchasing longitudinal completeness is purchasing the quality of that stitching, and there is currently no way to assess it other than by inspection of delivered data.
The reported figures describe speed and breadth, which are real and checkable. Neither speaks to correctness.
Ask how entity resolution works and what confidence it carries, how conflicting records between two providers are reconciled and which wins, what the extraction accuracy is for photographed paper records, and what evaluation exists against a source of truth.
Consent architecture is the product thesis here rather than a compliance layer bolted on, and it produces the strongest position available in real world data. Patients aggregate their own records under federal access rights, and separately choose whether to contribute de identified data to research, and the separation matters as much as the consent: assembling your own history and donating it to a research corpus are two different decisions, and a product that collapses them into one signature has obtained agreement to the first and inferred it for the second.
Data provenance is stated as direct from the patient with no tokenisation or third party guesswork, and the company contrasts this explicitly with brokers stitching together claims and hospital records without patient participation, which is the arrangement most of this index's real world data records rest on.
Consent obtained at the source is the strongest stewardship position available, because every downstream question about basis, scope and withdrawal has an answer that traces to a person who was actually asked. Two residuals a buyer or a patient should still establish, and they are the reason to ask rather than reasons to discount. No model or model family, hosting arrangement or sub processor list was located, so the chain beneath the consent is unenumerated. And nothing states what happens to an aggregated record if a patient stops using the service, or whether a research contribution can be withdrawn once made.
Commercial validation is real, with the company reporting customers among a majority of the top 10 to 15 global pharmaceutical companies, and one described application involved analyzing a safety signal for a breast cancer drug where the platform found a lower incidence of a suspected adverse event than previously believed.
However that account is vendor described without published methodology, and no peer reviewed validation of the dataset's completeness or representativeness was retrieved. Patient contributed cohorts also carry self selection characteristics that a research buyer must assess directly.
Consent architecture is the product thesis rather than a compliance layer. Patients aggregate their own records under the 21st Century Cures Act access rights and separately choose whether to contribute de identified data to research, and the company contrasts this explicitly with third party brokers stitching together claims and hospital records without patient participation. Data provenance is stated as direct from the patient with no tokenization or third party guesswork. Consent obtained at the source is the strongest stewardship position available in real world data.
The regime does not apply, the company knows it, and it publishes for the regime that does. That combination is rare enough to earn the grade.
Records reach this company because a patient connects their own portals and directs the disclosure to an application they control. Where an individual directs a covered entity to send records to a third party app, the recipient does not become a business associate and the health privacy rule does not follow the data. The company is neither a covered entity nor a business associate, and the absence of a business associate agreement is a correct consequence of the design rather than a gap.
What is unusual is that it publishes a separate consumer health data privacy notice addressed to the state consumer health data laws that do apply to this position, naming the rights those laws confer to access, delete and withdraw consent. It enumerates the sensitive categories it handles with real candour, including reproductive and sexual health, pregnancy status, gender affirming care, and photographs that may constitute biometric information. Those are precisely the categories those statutes exist to protect, and most companies operating outside the health privacy rule publish nothing equivalent.
One drafting point holds this at B and is worth raising with the company. The consent checkbox at account creation is labelled as agreement to the health privacy rules and the terms of use. A patient does not agree to the privacy rule; they authorise a disclosure. Labelling the consent that way at the exact moment the data leaves the rule's protection invites the belief that its protections travel with it. They do not.
No SOC 2, ISO 27001 or equivalent attestation was located and there is no trust centre. The strongest security statement found is application store copy describing robust, verified security systems, which names nothing and is verified by no one.
The consequence of the company's own structure makes this weigh more than the grade suggests. Because records arrive by patient directed disclosure rather than through a covered entity relationship, the health privacy rule does not reach them, and neither does its security rule. The two travel together. There is therefore no federal requirement here for the administrative, physical and technical safeguards a buyer would ordinarily assume apply to a company holding complete medical records. What governs is the company's own terms, consumer protection law, the health breach notification rule and the state consumer health data statutes.
What sits behind that gap is unusually complete. Not claims data, not a partial hospital extract, but assembled longitudinal records reported at up to twenty years, drawn from more than fifty thousand provider connections, including radiological images, laboratory results and self reported detail, for a population the company describes as concentrating on the sickest and most complex patients.
One item to check rather than assume: the application's store privacy disclosures include a section for data used to track users across other companies' applications and websites. Read what is listed there before drawing conclusions either way.
Ask whether any independent security examination exists, what encryption and access controls apply, and who inside the company can read an identified record.
A scoping determination, and the framework that does apply is one where this company's design is a genuine advantage.
The platform assembles and structures records. It does not diagnose, treat or make a recommendation about an individual's care, so no device pathway attaches and none is claimed. The consumer application is a personal health record tool, which is likewise outside device regulation.
What governs the product actually being sold is the framework for real world data and real world evidence used to support regulatory decisions. That framework turns on two properties: whether the data is relevant to the question, and whether it is reliable, meaning accurate, complete, traceable to its source and consistently collected. Provenance is central to it.
On that measure this company's design is stronger than the alternative it positions against. Records are obtained directly from the originating providers through patient authorised access rather than purchased, tokenised and probabilistically matched by an intermediary, so the chain from source system to analytic record is short and traceable in principle. That is a real regulatory asset and a sponsor should understand it as one.
What is not published is whether that potential has been realised: no documentation of data accrual and missingness, no validation of assembled records against source, no audit trail specification, and no statement of whether the dataset has supported any regulatory submission. Traceable in principle and traceable in an inspection are different things.
Ask for the data quality documentation, the missingness profile by data type and provider, and whether any submission has relied on this dataset.
No governance framework, model documentation or evaluation was located, and two domain questions apply with force. Both follow from what makes the product good.
The first is composition. This dataset is built from people who download a health application, connect their portals and consent to research. The company states plainly that it concentrates on patients with serious and complex conditions, because those patients carry the heaviest record keeping burden and gain most from the tool. That is a sound product strategy and it is also a selection mechanism. The resulting cohort will over represent people who are digitally capable, motivated by illness, insured enough to have generated records across many providers, and willing to contribute to research. The company positions this dataset against broker assembled claims data, correctly noting that broker data is fragmented. Broker data carries its own bias, driven by who is insured. Patient contributed data carries a different one, driven by who participates. A research buyer needs the second characterised as much as the first, and nothing published characterises it.
The second is record linkage. The stated capability is stitching fragmented sources into a coherent longitudinal journey, which is an inference task with real error modes: merging two people, splitting one, or resolving conflicting entries wrongly. No accuracy figure, validation method or error rate was located for any of it, and a linkage error propagates silently into every analysis built on that patient.
One effort points the right way and deserves credit: the company participates in a cross industry initiative examining disparities in breast cancer treatment adherence, which is exactly the kind of work this dataset is well placed to support. Ask for cohort composition against a reference population, and for linkage accuracy.
The functional description is clear and nothing beneath it is disclosed. Public material describes connection to tens of thousands of provider sources, automatic retrieval once portals are linked, paper documents accepted through a camera, and assembly into a structured longitudinal journey spanning up to twenty years in around thirty seconds. Those figures describe speed and breadth, which are real and checkable, and neither speaks to correctness.
What is absent is the layer that is the entire technical proposition. The differentiator is coherence across sources rather than volume, so a buyer purchasing longitudinal completeness is purchasing the quality of the stitching, and no account exists of how record linkage and entity resolution are performed, what confidence attaches, how conflicting entries between two providers are reconciled and which one wins, what extraction accuracy applies to photographed paper records, or what evaluation has been done against any source of truth.
The reconciliation question is the sharpest, because two providers disagreeing about a medication or an allergy is common, and a silent resolution rule produces a record that looks authoritative and settles a conflict the patient never saw. No warranty, indemnity or remediation commitment was located. Ask how entity resolution works and what it does when uncertain, the reconciliation rule, extraction accuracy on camera captured documents, and whether a user is shown that sources conflicted.
Retrieval breadth is the core competency: connections to more than 50,000 US healthcare providers spanning Epic, Oracle Health, athenahealth, Quest Diagnostics, and the VA, with paper record capture by camera as a fallback. This is patient mediated access under federal interoperability rules rather than institutional integration, a materially different and broader path than negotiating system by system.
Nothing was located on hosting location, region, tenancy, retention or subprocessors, and the architecture creates a specific question that is not addressed.
What is visible: a consumer application distributed through both major mobile stores alongside a web platform, connecting to more than fifty thousand provider sources through patient authorised portal connections, operating in the United States only. Single country operation removes most of the cross border complexity that burdens other vendors in this lane, which is a genuine simplification.
The question the design raises is separation. There are two data estates here doing different jobs. One is the patient's own consolidated record, identified by construction, containing radiological images, laboratory results and self reported detail, which the patient reads and corrects. The other is the de identified research dataset licensed to pharmaceutical customers. Nothing published describes how those are separated, whether de identification produces a genuinely distinct store or a view over the same records, where either sits, or what a pharmaceutical customer's access actually reaches.
The imaging component deserves its own question. Radiological images are large binary objects that are difficult to de identify reliably, because identifiers persist in embedded metadata and, for some modalities, in the image itself.
One inference cuts in the company's favour. It commits to deleting a user's entire account and removing all records from its systems, which is only implementable if a complete data map exists. Ask to see it: where records are stored and processed, what the retention schedule is, how the identified and de identified estates are separated, and which subprocessors touch either.
The structure is clearly disclosed even though amounts are not: the patient application is free, and revenue comes from de identified data and insight products sold to pharmaceutical and diagnostics customers. A buyer understands exactly how the two sides relate, which is more than most data vendors disclose. No rate card published.
Precisely bounded on both sides of the business, and the company states both rather than blurring them.
On collection: United States patients drawing records from United States providers, spanning the major record platforms, a national laboratory network and the veterans health system, with paper capture as a fallback for what portals do not hold. The geography is explicit and the company does not claim reach it does not have.
On sale: biopharmaceutical teams in health economics and outcomes research, market access, clinical operations and medical affairs. Naming the functions rather than gesturing at pharmaceutical customers generally is useful, because those four buy different things and value different properties of a dataset. The reported therapeutic coverage spans oncology, rare disease and cardiometabolic conditions among others, which is consistent with a cohort concentrated on patients with serious and complex illness.
The scoping that matters most is the one the company draws itself, between the free consumer application and the data product sold to industry. It is explicit that the application is the collection mechanism and the licensed data is the product, and that patients choose separately whether to contribute. A reader can therefore tell exactly what is being sold to whom, which is not true of every company in this part of the market.
A buyer should note only that coverage of a therapeutic area is not the same as depth within it, and should ask for cohort counts by condition before assuming a rare disease population is large enough to support an analysis.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Contact the vendor
|
De identified data and insight products sold to life sciences; patient application free | — | — | Vendor Published |
Two sided by design and clearly disclosed: the patient facing application is free, and revenue comes from de identified, patient consented data and insight products sold to pharmaceutical and diagnostics organizations. No rate card published for the data products.