Turning Real Medical Conversations into AI Training Data

Turning Real Medical Conversations into AI Training Data

The Data Hiding in Every Doctor’s Visit

There’s a version of every doctor’s visit that never makes it into the medical record. The way a patient struggles to describe a symptom. The follow-up question that changes the conversation. The pause before a doctor finally says what they think is wrong. Most of that disappears once the visit is reduced to a clinical note.

And that’s exactly what makes the conversation so valuable for AI. A note is the finished version of the visit. the conversation is everything that happened along the way. But capturing that conversation isn’t as simple as pressing record. Accents, background noise, interruptions, medical terms and even small sounds like “mm-hm” can change what a system understands. In one study of primary-care encounters, overall speech-recognition error was around 12%, while some clinically relevant conversational sounds had error rates above 90% [1].

So the real challenge isn’t just collecting more hours of audio. It’s collecting the right conversations, transcribing them accurately, understanding what’s being said, and turning that information into something a model can actually learn from, while making sure every conversation was properly consented, licensed and de-identified along the way. That’s where the real value of clinical conversation data begins.

Where This Data Goes

Once you have enough real conversations, the possibilities go well beyond simply turning speech into text. One of the clearest examples is the ambient clinical scribe: software that listens during a consultation and turns the conversation into a clinical note, giving the doctor more time to focus on the person sitting in front of them. The reason this has become such an important area of clinical AI is straightforward: creating notes from doctor-patient encounters is time-consuming, while the conversations themselves are rarely recorded and are difficult to share because of their sensitive nature [5].

But the transcript is only the beginning. The same conversation can be used to identify symptoms, diagnoses, medications, procedures and other clinical details buried in ordinary speech. Over time, those structured pieces of information can become useful training data for systems that do more than document a visit. They can help organize clinical information and support a doctor as they work through a case [5].

Figure 1: From a real conversation to what it becomes: a clinical note, structured data, and eventually, diagnostic support.

Research is already moving in that direction. Google’s AMIE system, for example, was evaluated across 159 simulated clinical scenarios sourced from India, Canada and the UK, with scenarios spanning 51 medical specialties and 168 medical conditions or visit reasons. The study looked beyond diagnosis alone, evaluating history-taking, management, communication and other aspects of the clinical interaction [2]. The point isn’t that AI has replaced the doctor. It hasn’t. The more interesting point is that researchers are teaching AI to understand the conversation around medicine, not just the medical text that comes after it.

Study Scale Key finding
Singapore, 2026 [3] 169 consultations 15% less documentation time, 10.6% more eye contact
JAMA, 2026 [4] 8,581 clinicians 16 min less documentation time, 13.4 min less EHR time
JAMA, 2026 1,809 AI-scribe adopters 0.49 more visits/week

Figure 2: Measured impact of ambient AI in clinical practice.

Perit’s Role in This

This is where Perit comes in. The goal isn’t simply to collect recordings. It’s to turn real clinical conversations into data that AI systems can actually learn from.

Perit has built a catalog of more than 300,000 hours of real doctor-patient conversations, covering multiple specialties, geographies and care settings. The conversations come from both telehealth and in-clinic encounters, and are collected with the necessary consent and licensing before being de-identified to remove personally identifiable information.

What makes that collection useful is the structure around the audio. A raw conversation can be transcribed, separated by speaker, and annotated for the clinical information inside it, from symptoms and diseases to medications, diagnoses and other medical concepts. That turns hours of spoken conversation into a dataset with context, labels and structure, rather than just a folder full of recordings.

This is the difference between having medical audio and having medical training data.

Figure 3: From clinical conversations to AI-ready data.

What Actually Makes Data Usable

A large dataset sounds impressive, but volume alone doesn’t make it useful. For clinical AI, the value comes from several things working together:

Usable Clinical Data = Diversity × Quality × Annotation × Context

Diversity means covering different specialties, accents, patient populations, geographies and care settings. Quality means clear recordings, accurate transcription and reliable speaker identification. Annotation adds clinical meaning by labeling things such as symptoms, diseases, medications, diagnoses and procedures. And context connects those pieces to the conversation they came from, including who said what, why it was said and what happened during the encounter.

If any one of these is weak, the dataset becomes harder to use. A huge collection of perfectly recorded conversations with poor annotation is still limited. Likewise, beautifully labeled data from only one narrow type of consultation won’t teach a model how medicine actually sounds across different situations.

That is why building useful clinical training data is not just a question of scale. It is a question of how much information survives the journey from conversation to dataset.

Figure 4: Less documentation, more time focused on the patient. [3]

Beyond Audio: Imaging Data

Audio captures what happens in the conversation. But a large part of modern medicine happens in the images around that conversation.

A CT scan or X-ray is not simply a picture sitting in a folder. Medical imaging is usually stored in DICOM, a standard that carries both the image and information about the study around it. A single imaging study can contain multiple series and multiple images. A CT study can contain hundreds of individual slices. The same study can also be linked to a radiology report describing what the radiologist saw [6].

That structure matters for AI. An image on its own tells only part of the story. The associated metadata can provide information about the study and acquisition, while the radiology report provides a clinical interpretation of the findings. Public datasets demonstrate the value of keeping these pieces connected: MIMIC-CXR, for example, contains 227,835 imaging studies and 377,110 chest X-ray images, with each study linked to a free-text radiology report [6].

And the information in a report can be much richer than a simple diagnosis label. Radiology AI research has moved toward annotating anatomy, observations, uncertainty and relationships between findings. The RadGraph dataset, for example, was created using radiologist annotations to identify entities such as anatomy and observations and relationships such as where an observation is located. Later work has expanded this approach to CT, MRI and X-ray reports, showing how imaging data can be structured at a much finer level than simply “normal” or “abnormal.” [7] [8]

Figure 5: An imaging study is more than the image itself.

There is another challenge that is easy to overlook: privacy doesn’t stop at the edge of the image. DICOM files contain metadata that can include identifying information, and patient information can sometimes also be embedded directly into the pixels or graphics of an image. Proper de-identification therefore has to consider both the metadata and the image itself. The DICOM standard specifically addresses approaches for protecting identifying attributes and cleaning identifying information that may be burned into pixel data [9].

But de-identification is not simply about deleting everything. Some information is valuable precisely because it provides context. Study identifiers, acquisition information, modality, anatomy and relationships between multiple studies can all help preserve the usefulness of the dataset. The challenge is to remove what can identify a patient without removing the information an AI system needs to understand the case [9] [10].

That is why imaging data preparation is a balancing act between privacy, structure and clinical usefulness.

Figure 6: Preparing medical imaging for AI.

The result is not just a collection of de-identified scans. It is a structured imaging dataset in which the image, its clinical context and its annotations remain connected. That structure also makes it possible to look beyond a single scan. When multiple studies from the same patient can be linked over time, AI systems can begin to learn from change, including what appeared before, what changed and what stayed the same.

The Challenges of Building Clinical AI Data

But turning all of this into useful AI data is not straightforward. Medical conversations and imaging files carry sensitive patient information, so privacy has to be considered from the very beginning. Names, diagnoses and medications can appear in conversations, while DICOM files can contain identifying metadata or information within the image itself. The challenge is to remove what could identify a patient without removing the clinical context that makes the data valuable.

There is also the problem of variation. Doctors and patients speak differently across specialties, regions and care settings. Accents, languages, background noise, interruptions and medical terminology can all affect how a conversation is transcribed and understood. Clinical annotation adds another layer because meaning often depends on context. A doctor mentioning a disease does not necessarily mean the patient has that disease, just as a finding in an imaging report can carry uncertainty or depend on what else appears in the study.

And then comes scale. Processing a small number of conversations or scans is very different from doing it across hundreds of thousands of them while keeping transcription, annotation, quality checks and privacy consistent. The goal, ultimately, is not just to collect more data. It is to make sure that the clinical story within that data stays accurate, connected and useful.

The Future of Medicine

The real opportunity isn’t in teaching AI to understand one doctor’s visit or one medical image. It’s in helping it understand the story that connects them.

Imagine a patient who visits a doctor with recurring symptoms. The first consultation captures what the patient describes and what the doctor suspects. A few weeks later, a CT scan adds another piece of evidence. Months later, a follow-up consultation reveals that the symptoms have changed. Another scan shows what has changed or hasn’t.

Today, much of that information is stored across different systems, formats and moments in time. A conversation sits in one place, an imaging study somewhere else, and a clinical report somewhere in between. For an AI system, bringing those pieces together is a very different challenge from simply reading a single note.

That’s where multimodal and longitudinal clinical AI becomes interesting. Recent research is already moving beyond text-only conversations, testing systems that can reason across conversations alongside clinical documents, images, ECGs and other medical information. In one 2026 Nature Medicine study, a multimodal version of Google’s AMIE was evaluated on simulated telehealth consultations that included clinical documents, skin photographs and ECGs. The system was assessed not only on diagnosis, but on how well it incorporated these different sources of evidence into its reasoning [11].

The longer-term possibility is even more powerful: AI that can understand change over time. Instead of treating every appointment as an isolated event, a system could potentially connect what was said during one visit with what was seen in an earlier scan, what medication was prescribed later, and what happened at the next follow-up.

That doesn’t mean AI replaces the doctor. The more realistic goal is almost the opposite, giving doctors better access to the information they already need, without making them piece the patient’s history together manually every time.

And that brings the story back to the data itself. Better clinical AI starts long before a model is trained. It starts with capturing real conversations, preserving their context, structuring the clinical information inside them, and connecting that information with other forms of medical data.

The future of medical AI may depend less on how much data we collect, and more on how well we connect the data we already have.

References

  1. Mm-hm, Uh-uh: Are Non-Lexical Conversational Sounds Deal Breakers for Ambient Clinical Documentation Technology? JMIA, 2023.
    https://pmc.ncbi.nlm.nih.gov/articles/PMC10018260/.com
  2. Towards Conversational Diagnostic Artificial Intelligence, Nature, 2025.
    https://www.nature.com/articles/s41586-025-08866-7.com
  3. Impact of an Ambient AI Scribe Among Clinicians and Patients, JMIR, 2026.
    https://medinform.jmir.org/2026/1/e85580.com
  4. Changes in Clinician Time Expenditure and Visit Quantity With Adoption of AI-Powered Scribes, JAMA, 2026
    https://jamanetwork.com/journals/jama/article-abstract/2847319.com
  5. Aci-bench: A Novel Ambient Clinical Intelligence Dataset, Scientific Data, 2023.
    https://pmc.ncbi.nlm.nih.gov/articles/PMC10482860/.com
  6. MIMIC-CXR: A De-identified Database of Chest Radiographs With Free-Text Reports, Scientific Data, 2019.
    https://pmc.ncbi.nlm.nih.gov/articles/PMC6908718/.com
  7. RadGraph: Extracting Clinical Entities and Relations From Radiology Reports, PhysioNet.
    https://www.physionet.org/content/radgraph/1.0.0/.com
  8. RadGraph-XL, ACL Findings, 2024
    https://aclanthology.org/2024.findings-acl.765/.com
  9. DICOM PS3.15: Security and System Management Profiles, DICOM Standard.
    https://dicom.nema.org/medical/dicom/current/output/chtml/part15/chapter_E.html.com
  10. Guidance on Methods for De-identification of Protected Health Information, U.S. Department of Health and Human Services.
    https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html.com
  11. Advancing Conversational Diagnostic AI With Multimodal Reasoning, Nature Medicine, 2026.
    https://www.nature.com/articles/s41591-026-04371-0.com

 

Was this article helpful? 4/5 1 rating