What happened to patients who were already treated — the procedure, the medications, the side effects, the course afterward. It is written down, in narrative, across encounters, in records nobody has time to read at scale. We read them, and we show the passage behind every finding.
Retrospective, not predictive. Evidence, not discovery.
Structured fields capture what is billable. The clinically interesting part — how the patient responded, what was tolerated, what was stopped and why, what the physician noticed and wrote down — lives in narrative text. Reading a chart properly takes most of an hour, which is why it is done on a sample of high-value records and not on the rest.
Comprehensive chart review is measured in tens of minutes per record. That cost sets the ceiling on how many records get looked at.
Two competent reviewers reading the same chart do not always extract the same facts. Inter-rater variability is a known limitation of manual abstraction.
Under time pressure, information buried deep in a long record is missed — and the buried part is often the part that mattered.
What a health system learns about how a therapy performed in practice rarely reaches the people who made it, in any usable, structured form.
The extraction and de-identification layer this division needs is the same layer our Revenue Integrity work already required — because the objection that shaped that work was "the foundation of this starts before the data set you are looking at." Both programs needed to read the record rather than the form. That shared substrate is the reason these two very different divisions exist in one program at all.
This division is in development. What follows is what it is designed to answer, stated as design intent rather than as a shipped feature list.
Better ways to treat and diagnose. Across a population you already serve: which treatment pathways were followed, what varied between them, and what the documented course looked like afterward. Where a diagnosis was reached late, what appeared in the record before it.
Structured evidence out of unstructured records. Cohort assembly against real inclusion criteria, with abstraction that shows its work — each extracted variable traceable to the passage it came from, so a reviewer can verify rather than trust.
Real-world usage feedback. De-identified, aggregated signal on how a product was actually used and tolerated in practice — dosing as administered rather than as labeled, concomitant therapy, documented side effects, and the reasons treatment was changed or stopped.
Anything returned to a manufacturer is de-identified and aggregated at the population level, under the customer's own governance and data-use terms, and never as patient-level records. Signal that could constitute a reportable adverse event is a human escalation path, not an automated output — the system routes it to a qualified reviewer and does not decide for them. Pharmacovigilance reporting is a regulated obligation belonging to the parties who hold it, and we do not discharge it on anyone's behalf.
Steps 1 and 2 are shared with Division One. Everything from step 3 onward is specific to this division and runs against a separate corpus with separate access control.
An extracted variable is stored with the passage it was derived from and the version of the extraction logic that produced it. A reviewer can open any value and read the sentence behind it. Without that, a real-world evidence dataset is an assertion; with it, it is reviewable.
The system is explicit about what it may conclude alone and what it must hand to a person — low-confidence extraction, contradictory documentation, and anything with a safety dimension. The escalation rules are configuration, not model behavior, so they can be reviewed and audited.
This division works exclusively on records that already exist, about care that has already been delivered. That single sentence draws most of the boundary.
| Dimension | What we do | What we do not do |
|---|---|---|
| Input | Records already collected in the course of care — notes, encounters, medications, procedures | No biological samples, no assays, no laboratory work of any kind |
| Time direction | Retrospective. What was documented and what followed | No prediction of whether an individual patient will respond to a therapy |
| Scientific claim | Signal and association, surfaced with its evidence, for qualified people to interpret | No causal claim, no biomarker discovery, no target identification, no mechanism |
| Regulatory posture | Analytics under the customer's own governance and data-use terms | Not a medical device, not FDA-cleared, not a diagnostic, not clinically validated |
| Clinical role | Input to review by clinicians, researchers, and medical affairs professionals | Never a treatment recommendation, and never a substitute for clinical judgment |
| Manufacturer output | De-identified, population-level, under the customer's terms | No patient-level data, no re-identifiable output, no assumption of anyone's pharmacovigilance obligation |
Organizations doing genuine discovery science — generating novel measurements in a lab and building models on them — need a records layer that is HIPAA-boundaried, auditable, and citable, and they generally would rather not build one. That is the work we do. Different question, different input, different part of the value chain.