Clinical Evidence
What happened to patients who were already treated — the procedure, the medications, the side effects, the course afterward. It is written down, in narrative, across encounters, in records nobody has time to read at scale. We read them, and we show the passage behind every finding.
Retrospective, not predictive. Evidence, not discovery.
The most underused clinical dataset is the one already in the chart.
Structured fields capture what is billable. The clinically interesting part — how the patient responded, what was tolerated, what was stopped and why, what the physician noticed and wrote down — lives in narrative text. Reading a chart properly takes most of an hour, which is why it is done on a sample of high-value records and not on the rest.
Manual
Comprehensive chart review is measured in tens of minutes per record. That cost sets the ceiling on how many records get looked at.
Inconsistent
Two competent reviewers reading the same chart do not always extract the same facts. Inter-rater variability is a known limitation of manual abstraction.
Incomplete
Under time pressure, information buried deep in a long record is missed — and the buried part is often the part that mattered.
Unshared
What a health system learns about how a therapy performed in practice rarely reaches the people who made it, in any usable, structured form.
Why we can build this and did not have to start over
The extraction and de-identification layer this division needs is the same layer our Revenue Integrity work already required — because the objection that shaped that work was "the foundation of this starts before the data set you are looking at." Both programs needed to read the record rather than the form. That shared substrate is the reason these two very different divisions exist in one program at all.
Three audiences, three different questions.
This division is in development. What follows is what it is designed to answer, stated as design intent rather than as a shipped feature list.
For clinicians and quality teams
Better ways to treat and diagnose. Across a population you already serve: which treatment pathways were followed, what varied between them, and what the documented course looked like afterward. Where a diagnosis was reached late, what appeared in the record before it.
- Care-pathway variation surfaced from the notes, not from claims alone
- Documented outcomes and complications assembled per cohort
- Earlier-signal review: what was present in the record before the diagnosis was made
For research organizations
Structured evidence out of unstructured records. Cohort assembly against real inclusion criteria, with abstraction that shows its work — each extracted variable traceable to the passage it came from, so a reviewer can verify rather than trust.
- Cohort identification from narrative, not just coded fields
- Variable-level provenance for every abstracted value
- Consistent extraction across reviewers and across time
For manufacturers
Real-world usage feedback. De-identified, aggregated signal on how a product was actually used and tolerated in practice — dosing as administered rather than as labeled, concomitant therapy, documented side effects, and the reasons treatment was changed or stopped.
- Usage patterns in practice versus label
- Documented adverse experience and tolerability signal
- Discontinuation and switching, with the stated reason where the record gives one
On manufacturer reporting specifically
Anything returned to a manufacturer is de-identified and aggregated at the population level, under the customer's own governance and data-use terms, and never as patient-level records. Signal that could constitute a reportable adverse event is a human escalation path, not an automated output — the system routes it to a qualified reviewer and does not decide for them. Pharmacovigilance reporting is a regulated obligation belonging to the parties who hold it, and we do not discharge it on anyone's behalf.
Assembled from the record, grounded in the record.
Steps 1 and 2 are shared with Division One. Everything from step 3 onward is specific to this division and runs against a separate corpus with separate access control.
What "grounded" means here
An extracted variable is stored with the passage it was derived from and the version of the extraction logic that produced it. A reviewer can open any value and read the sentence behind it. Without that, a real-world evidence dataset is an assertion; with it, it is reviewable.
What "escalation" means here
The system is explicit about what it may conclude alone and what it must hand to a person — low-confidence extraction, contradictory documentation, and anything with a safety dimension. The escalation rules are configuration, not model behavior, so they can be reviewed and audited.
What this is, and firmly what it is not.
This division works exclusively on records that already exist, about care that has already been delivered. That single sentence draws most of the boundary.
| Dimension | What we do | What we do not do |
|---|---|---|
| Input | Records already collected in the course of care — notes, encounters, medications, procedures | No biological samples, no assays, no laboratory work of any kind |
| Time direction | Retrospective. What was documented and what followed | No prediction of whether an individual patient will respond to a therapy |
| Scientific claim | Signal and association, surfaced with its evidence, for qualified people to interpret | No causal claim, no biomarker discovery, no target identification, no mechanism |
| Regulatory posture | Analytics under the customer's own governance and data-use terms | Not a medical device, not FDA-cleared, not a diagnostic, not clinically validated |
| Clinical role | Input to review by clinicians, researchers, and medical affairs professionals | Never a treatment recommendation, and never a substitute for clinical judgment |
| Manufacturer output | De-identified, population-level, under the customer's terms | No patient-level data, no re-identifiable output, no assumption of anyone's pharmacovigilance obligation |
Stated plainly, because the distinction is easy to blur
- We are not a life-sciences discovery company. We do not do proteomics, genomics, or any wet-lab science, and we never will
- We do not predict which patient will respond to a drug. Companies that do that work from laboratory measurements we do not generate and would not attempt to
- Our contribution is engineering — the governed, auditable data layer between records and analysis. It sits underneath scientific work; it is not scientific work
- This division is in development. There is no published result and no validation study yet
We are the layer under the science, not the source of it.
Organizations doing genuine discovery science — generating novel measurements in a lab and building models on them — need a records layer that is HIPAA-boundaried, auditable, and citable, and they generally would rather not build one. That is the work we do. Different question, different input, different part of the value chain.