Targeted narrative review · Published evidence
WoundScribe Evidence Review
An assessment of the published evidence associated with WoundScribe, including what the studies support, what they do not establish, and what health systems should verify before clinical evaluation.
Why the evidence matters
The Institute for Precision Health exists to help organisations adopt health science and technology well. In digital wound care that mandate has a specific shape, because the category is unusually easy to enter and unusually hard to substantiate. Assessment leans on visual judgment, which varies between clinicians and between visits. Progress is meaningful only across time. And one encounter must satisfy clinical, compliance, and reimbursement requirements simultaneously.
The consequence is that product demonstrations alone are insufficient. Published methods make it possible to examine study design, endpoints, limitations, funding, and conflicts of interest. The appropriate first question is therefore not simply whether a publication exists, but whether its methods and outcomes directly support the capabilities and clinical uses being considered.
How the evidence was assessed
This targeted narrative review assesses published evidence rather than testing the product directly. It is not a systematic review or meta-analysis. The method is stated so readers can understand how the evidence was located, classified, and interpreted, and can judge how much weight to place on the conclusions.
- Review questionWhat published evidence is directly or indirectly associated with WoundScribe, which platform components were studied, and what conclusions are justified by the reported methods and outcomes?
- Information sourcesA targeted search of PubMed and publisher records was conducted on 23 July 2026, supplemented by reference and citation checking and review of WoundScribe’s public science and product pages.
- Search conceptsBrand and product terms—WoundScribe, WoundScribe AI, and Pingoo AI—were combined with diabetes care, diabetic foot, limb preservation, ambient clinical documentation, patient education, medical knowledge extraction, retrieval-augmented generation, and the names of identified study authors.
- SelectionTwo journal articles met the inclusion criteria. Studies were eligible when they evaluated WoundScribe by name or examined a method publicly represented as part of the platform. Company material was used only to describe stated capabilities; testimonials and marketing statements were not treated as scientific evidence.
- Critical appraisalEach study was considered for directness to the deployed product, evaluation setting, comparator, sample and outcome definition, external validation, independence, author conflicts, funding, and applicability to clinical use. Source provenance and study quality were treated as separate questions.
- LimitsThe search was targeted rather than exhaustive and was not designed to identify every potentially relevant record. No product testing, site visit, data audit, code review, or independent clinician interview was conducted. Only two heterogeneous studies met the inclusion criteria, so statistical combination was neither appropriate nor attempted.
The published science
Two published studies associated with the developer were identified. They examine different parts of the technical approach and should not be treated as evidence for components they did not test.
Study 1: retrieval-grounded medical knowledge and patient education
A study in the Journal of Diabetes Science and Technology, indexed in PubMed under PMID 38767382, reports a retrieval-augmented generation model intended to answer questions about diabetes and diabetic-foot care for lay readers (reference 1). The knowledge base was assembled from 295 articles, and the model was tested using 175 questions. The authors used zero-shot and few-shot prompting, expert assessment, and a true-or-false evaluation to examine factuality and comprehension.
The study reported 98% accuracy under its evaluation procedure. That result supports the feasibility of a specific medical-knowledge-extraction and patient-education approach under the reported test conditions. Some reference labels were revised after discrepancies between human labels and model reasoning were examined, so the reported accuracy should be interpreted as a post-adjudication result rather than an untouched external-validation result. The study does not evaluate wound imaging, measurement, coding, healing analytics, real-world workflow performance, patient outcomes, or the complete commercial platform.
The study was sponsored by Silverberry Group. Its disclosures identify Silverberry relationships among several authors, including officer/shareholder, employee, and research-student roles. Peer review makes the methods and findings available for scientific scrutiny, but does not by itself establish reproducibility, independence, clinical effectiveness, or applicability to every platform component.
Study 2: ambient documentation and multi-agent patient education
A 2026 study in the Journal of the American Podiatric Medical Association evaluates WoundScribeAI’s ambient documentation and Pingoo patient-education components by name (reference 2). It reports two phases: 30 prerecorded physician–patient conversations drawn from training videos and 30 simulated consultations conducted by clinicians. No actual patients participated. The system generated transcripts, SOAP notes, visit summaries, and conversational educational content, which reviewers assessed using five-category rating scales.
For SOAP-note accuracy, 86.7% of evaluated outputs were rated “Excellent” and 13.3% “Very Good.” For patient-summary accuracy, 90% were rated “Excellent” and 10% “Very Good.” All evaluated SOAP notes and patient summaries were rated as having no hallucination. These are descriptive ratings from the study’s controlled and simulated sample, not error rates measured in routine clinical deployment. The paper reports no blinded independent external evaluation, no control group for patient or workflow outcomes, and no inferential comparison establishing clinical benefit.
This study is more direct to the commercial product than the earlier knowledge-extraction paper and provides preliminary support for documentation and educational-content generation under the conditions tested. It does not constitute a prospective study in routine clinical care and does not establish performance for imaging, wound measurement, coding, or healing analytics. Reports of no hallucinations apply only to the evaluated outputs and should not be generalized to other inputs, populations, deployments, or future model versions. The paper also states that its underlying data are not publicly available because of confidentiality restrictions, which limits independent reanalysis.
This study was also sponsored by Silverberry Group and includes authors with disclosed relationships to the sponsor. Those relationships do not invalidate the work, but they increase the importance of replication and prospective evaluation by investigators independent of the developer.
Overall appraisal of the published evidence
Together, the studies provide early evidence for selected knowledge-extraction, documentation, and patient-education functions. Both involved the developer and used technical or simulated evaluations. Neither establishes the safety, effectiveness, generalisability, or clinical utility of the complete WoundScribe platform in routine care. A structured pilot remains the appropriate next step.
What the platform does
The company describes an AI-native platform that carries a wound from photograph through to clinical documentation, with six capability areas working as one workflow: wound imaging and measurement, ambient clinical scribing, structured charting, reimbursement coding, healing analytics, and patient education. It states that the platform serves providers across clinics, hospitals, and mobile or home-based settings, that it supports integration with existing EMR systems, and that it takes a HIPAA-compliant, privacy-by-design approach.
The positioning is explicitly against single-purpose tools, on the argument that fragmented systems can require clinicians to re-enter information across multiple workflows. The integrated-workflow rationale is plausible and consistent with broader concerns about documentation burden and fragmented health-information systems. However, the operational benefit, interoperability, reliability, and total correction burden of WoundScribe’s implementation have not yet been independently established.
These are company-stated capabilities, reproduced here because a review should describe what it is reviewing, and labelled as such because scope claims are properly settled in a pilot rather than in a review.
Evidence standing at a glance
with developer involvement
A PubMed-indexed study of retrieval-grounded diabetes and diabetic-foot education (reference 1), and a study of WoundScribeAI ambient documentation and patient education using prerecorded and simulated encounters (reference 2). Both studies received Silverberry sponsorship and include authors with disclosed sponsor relationships.
capabilities
Platform scope and the six capability areas, care settings served, EMR integration, and privacy posture (references 3 and 4). These statements define claims that should be verified in a pilot and in technical, privacy, and contractual documentation.
in evaluation
Prospective clinical outcomes, independently measured time savings, coding accuracy, wound-measurement accuracy against a reference standard, subgroup performance, deployment reliability, and performance of the complete integrated workflow. No claim is made about these outcomes in either direction.
What to verify before adoption
None of what follows is particular to this platform. It is the standard the Institute for Precision Health teaches for any clinical AI system, applied here as it would be to any other.
- Correspondence between publication and productAsk which published components are in production today and what has changed since publication.
- Provenance of every performance numberMeasured on what population, by whom, against what comparator, and is the underlying study readable?
- Measurement validity across skin tones and wound typesRequest wound-specific performance stratified by skin tone, wound type, care setting, device, lighting, distance, and image-acquisition conditions. Broader clinical-AI literature shows that algorithmic performance can vary across populations and image-based contexts (references 7 and 8), but wound-specific performance must be demonstrated directly.
- The human-oversight designMap where a clinician reviews, edits, and signs, and confirm that correcting the system is as easy as accepting it. Human-factors research shows why automation bias and complacency must be considered in oversight design (references 5 and 6).
- Error behaviour, not just accuracyHow the system behaves when uncertain—whether it abstains, flags, or asserts—matters as much as how it behaves when correct.
- Data governance and privacy documentationData location, retention, de-identification, subprocessors, and whether customer data trains models — settled in contract.
- Workflow fit under real conditionsTime the task end to end with your own clinicians, including correction and exception paths.
- Monitoring and rollbackDefine what is measured after go-live, who owns it, and what triggers rollback.
The precision-health perspective
The question that matters most in this category is not only whether a product works, but whether adopting it leaves an organisation with better data than it had before. Precision health depends on data that is structured, comparable across time and sites, and captured without introducing new bias — and the data model a platform imposes in year one determines what can be proven in year three.
An integrated approach could offer a scientific advantage over point tooling if it captures wound assessment, documentation, and progression using valid, consistent, interoperable definitions. Under those conditions, it could support a more longitudinal and comparable record. Integration alone does not guarantee data quality and can propagate systematic error across a workflow, so measurement validity, completeness, provenance, correction history, and exportability should be evaluated explicitly.
Conclusion
The identified literature provides preliminary evidence for selected components of WoundScribe’s approach, particularly retrieval-grounded patient education, ambient transcription, SOAP-note generation, visit summaries, and educational-content production. Both studies involved the developer and were conducted using controlled materials or simulated consultations. They do not establish the clinical effectiveness, safety, real-world workflow benefit, subgroup performance, or patient outcomes of the complete deployed platform.
The evidence is sufficient to justify consideration of a structured, independently governed pilot, but not to conclude that the full platform is clinically effective. Such a pilot should collect baseline measurements before deployment, specify comparators and clinically relevant endpoints, evaluate correction burden and subgroup performance, document model and software versions, and use pre-agreed monitoring and rollback triggers. Organisations considering evaluation can review WoundScribe’s published material and science and apply the questions above.
Disclosure and references
Disclosure. The Institute for Precision Health and WoundScribe are affiliated through Silverberry Group. This assessment is not independent. No separate fee was charged for producing this article. References 1 and 2 report Silverberry sponsorship and include authors with disclosed relationships to the sponsor. Those relationships are relevant when interpreting the findings and increase the importance of independent replication.
Published research
- Mashatian S, Armstrong DG, Ritter A, Robbins J, Aziz S, Alenabi I, Huo M, Anand T, Tavakolian K. Building Trustworthy Generative Artificial Intelligence for Diabetes Care and Limb Preservation: A Medical Knowledge Extraction Case. J Diabetes Sci Technol. 2025;19(5):1264–1270. doi:10.1177/19322968241253568. PMID 38767382. PubMed record — accessed 23 July 2026. The paper reports Silverberry sponsorship and Silverberry relationships among several authors.
- Mashatian S, Wung SF, Ritter A, Fishman J, Robbins J, Aziz S, Huo M, Armstrong DG. Responsible AI for Personalized Patient Education and Engagement Across Medical Conditions: Leveraging Multi-Agent LLMs, Ambient Technology, and NotebookLM—A Case Study in Diabetes Education and Limb Preservation. J Am Podiatr Med Assoc. 2026;116(3):30. doi:10.3390/japma116030030. Publisher record — accessed 23 July 2026. The paper reports Silverberry sponsorship and relationships with the sponsor.
Independent published evaluation of the deployed product
No independent published evaluation of the deployed product was identified in the targeted search conducted on 23 July 2026. Because the search was not systematic, this statement should not be interpreted as proof that no such publication exists.
Company-stated product capabilities
- WoundScribe. Home. woundscribe.ai — accessed 23 July 2026.
- WoundScribe. Our Science. woundscribe.ai/our-science — accessed 23 July 2026.
Supporting literature, general to clinical AI
- Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. J Am Med Inform Assoc. 2012;19(1):121–127. doi:10.1136/amiajnl-2011-000089. PMID 21685142.
- Parasuraman R, Manzey DH. Complacency and bias in human use of automation: an attentional integration. Hum Factors. 2010;52(3):381–410. doi:10.1177/0018720810376055. PMID 21077562.
- Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447–453. doi:10.1126/science.aax2342. PMID 31649194.
- Daneshjou R, Vodrahalli K, Novoa RA, et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Sci Adv. 2022;8(31):eabq6147. doi:10.1126/sciadv.abq6147. PMID 35921452.
Not independently verified
Clinical outcomes, time savings, coding accuracy, denial rates, measurement accuracy, subgroup performance, pricing, deployment scale, and regulatory classification. No supporting source for these outcomes was identified in the targeted review.