As the adoption of AI ambient scribing technology continues to rapidly advance, the healthcare industry’s standard method for evaluating clinical AI notes may be fundamentally flawed and fail to detect critical errors or omissions in patient notes.
That’s the view of Suki researchers, based on their study of assessment rubrics. Suki provides ambient clinical artificial intelligence solutions for healthcare providers and works with over 400 healthcare systems.
Suki researchers argue that the Physician Documentation Quality Instrument (PDQI-9), the medical industry’s standard tool for evaluating AI-generated clinical notes developed in 2012, is poorly suited for evaluating modern ambient AI scribes. That tool, the PDQI-9, was originally validated on a very small sample of inpatient notes and focuses on overall note quality (such as organization and brevity) rather than identifying LLM-specific errors, the researchers wrote in a white paper.
The core assessment domains of the PDQI-9 are current, accurate, thorough, useful, organized, concise, integrated, internally consistent, and consistent.
But the basic quality rubric “needs a makeover in the post-LLM ambient world,” said Kevin Wang, MD, Suki’s chief medical officer.
Suki’s white paper analyzes the limitations of comprehensive Likert-based scoring tools, such as the PDQI-9, in reliably evaluating AI-generated clinical notes and detecting certain errors such as hallucinations. Suki provided initial content for the whitepaper to Fierce Healthcare.
In a study of 84 pairs of notes across four disciplines, Suki researchers found significant inconsistencies among reviewers, including disagreements over what constitutes a hallucination, unreliable ratings of accuracy, and variations in scores depending on the AI model that generated the note. They argue that these problems reflect a fundamental mismatch between traditional note quality rubrics and the types of errors that LLMs make.
The PDQI-9 considers factors such as organization, brevity, and cohesion to give an overall score, which the Suki researchers argue is precisely the wrong lens for LLM errors. Traditional assessment frameworks may not reliably detect more serious errors, such as fabricating drug doses, missing diagnoses, and omitting clinical details.
The white paper raises concerns about how health systems evaluate AI scribes around them and suggests that the industry needs more rigorous ways to measure the quality of AI documentation and patient safety risks.
Suki’s research was driven by the belief that medical settings need more nuanced ways to assess the quality of clinical notes produced by environmental AI, Wang said. After reviewing prior research and industry standards, the company concluded that existing measures may no longer be sufficient for evaluating AI documents. Suki researchers argue that older note quality research and evaluation frameworks, including the PDQI-9, were developed before the era of ambient AI and were designed to evaluate EHR-based inpatient documentation rather than AI-generated clinical notes. The company argues that these measures do not adequately capture the factors most important to ambient AI, including factual accuracy, consistency, alignment with clinician intent, and physician acceptance.
Wang said the white paper tracks the evolution of the assessment of note quality and factual accuracy and lays the foundation for a larger study that Suki plans to publish later this year. The company is considering a new evaluation rubric that aims to address existing gaps and better measure the quality of AI-generated documents.
“It will show why existing frameworks have limitations and why we need something better to prove the true quality of ambient technology,” he said.
Suki researchers also considered alternative frameworks (PDSQI-9, SCRIBE, FActScore, VeriFact, CREOLA, etc.). Although new tools attempt to address the limitations, they still face challenges of low reliability, the researchers concluded. None of these tools combine sentence-level error detection, inter-rater reliability as a first-class metric, and statistical procedures for model release decisions, according to the white paper.
Researchers argue that valid assessment requires sentence-level error-specific metrics with validated interrater reliability and statistical testing procedures.
Wang outlined some of the potential patient safety risks that can occur if LLM errors go undetected in AI-generated patient visit records. “Tuberculosis has two meanings: latent tuberculosis and one meaning that the tuberculosis is no longer active. That’s very different from previously treated and cured tuberculosis. If I were a doctor talking to a patient, I wouldn’t speak all of these jargons. But imagine that a note prints out one of these two, and statistically it’s because of the difference. One says active tuberculosis, but it’s not active now. The other one. ‘You’re saying you don’t have it, it’s clinically very different and pharmacologically very different,’ he explained.
“As a patient, you want to be sure that your healthcare provider is outputting the correct notes for your current clinical condition, and that’s just one type of risk,” he said, adding that upstream issues with note quality and accuracy can have downstream impacts on reimbursement and coverage.
Differences in the quality of clinical records provided to insurance companies can mean the difference between a procedure like a colonoscopy that is coded as a screening colonoscopy (often fully covered by the insurance company) and a diagnostic procedure that requires the patient to pay coinsurance or a deductible.
Wang argues that there needs to be more transparency among AI scribe vendors about how they evaluate the accuracy and quality of their technology’s output.
“We’d love to see the statistics. We’ve seen other writing companies talk about hallucinations. We’d love to see the inter-rater reliability. We think these should not be black boxes,” Wang said.
He noted that the “stakes” will be in showing the quality behind the AI scribe’s technology.
“More and more of the partners we contract with are doing their own quality assessments. I think the healthcare industry is moving in the direction of having to convince themselves that what they’re buying is higher quality than what they don’t have. I think that’s going to be the new standard,” he said.
As the adoption of AI in the healthcare sector advances rapidly, assessing quality is essential, Wang said.
“In a year or two, there could be hundreds or even dozens of AI vendors in a single health system. AI in Electronic Health Records There may be vendors. There may be specialty pharmaceutical companies. There may be hardware. Imagine how many times quality is undervalued in all of this. “At the end of the day, it’s the doctors, the health care providers, the clinicians who treat the actual patients. So that’s what I’m working on. In the future, I hope to have a lot more conversations like this, a lot more clinical and academic conversations.”
He added, “I think we’re going to see this movement where AI is driving innovation and great user experiences and driving new quality and evaluation experiences. It’s not the most glamorous or sexy headline, but so much of healthcare is based on old and outdated things that need to be reinvented. We have an opportunity to change that going forward.”

