Skip to main content

Main menu

  • HOME
  • ARTICLES
    • Current Issue
    • Archives
    • Special Collections
    • Abstracts In Press
  • INFO FOR
    • Authors
    • Reviewers
    • Call For Papers
    • Subscribers
    • Advertisers
  • SUBMIT
    • Manuscript
    • Peer Review
  • ABOUT
    • The JABFM
    • The Editing Fellowship
    • Editorial Board
    • Indexing
  • CLASSIFIEDS
  • Other Publications
    • abfm

User menu

Search

  • Advanced search
American Board of Family Medicine
  • Other Publications
    • abfm
American Board of Family Medicine

American Board of Family Medicine

Advanced Search

  • HOME
  • ARTICLES
    • Current Issue
    • Archives
    • Special Collections
    • Abstracts In Press
  • INFO FOR
    • Authors
    • Reviewers
    • Call For Papers
    • Subscribers
    • Advertisers
  • SUBMIT
    • Manuscript
    • Peer Review
  • ABOUT
    • The JABFM
    • The Editing Fellowship
    • Editorial Board
    • Indexing
  • CLASSIFIEDS
  • JABFM on Bluesky
  • JABFM On Facebook
  • JABFM On Twitter
  • JABFM On YouTube
Article CommentaryCommentary

The Importance of Primary Care Subject Matter Experts: Output Quality in Large Language Models Prompt Engineering

Winston Liaw, Quang Hung Bui, Omolola Adepoju and Hien Van Nguyen
The Journal of the American Board of Family Medicine June 2026, 39 (1) 162476; DOI: https://doi.org/10.3122/jabfm.2025.250326R1
Winston Liaw
1 Department of Health Systems and Population Health Sciences, Tilman J. Fertitta Family College of Medicine University of Houston https://ror.org/048sx0r50
MD, MPH
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Quang Hung Bui
2 Department of Electrical and Computer Engineering, Cullen College of Engineering University of Houston https://ror.org/048sx0r50
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Omolola Adepoju
1 Department of Health Systems and Population Health Sciences, Tilman J. Fertitta Family College of Medicine University of Houston https://ror.org/048sx0r50
3 Humana Integrated Health Systems Sciences Institute University of Houston https://ror.org/048sx0r50
PhD
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Hien Van Nguyen
2 Department of Electrical and Computer Engineering, Cullen College of Engineering University of Houston https://ror.org/048sx0r50
PhD
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • Article
  • References
  • Info & Metrics
  • PDF
Loading

Abstract

Large language models (LLMs) offer new opportunities to support primary care by generating documentation, summarizing records, answering messages, and providing education. While LLMs show promise, their effectiveness depends on how well their outputs align with real-world clinical needs. This commentary describes one team’s development of Primary Care (PC) Navigator, a multimodal artificial intelligence (AI) tool that combines audio and video from clinical encounters with an LLM to generate behavior change plans. Our pilot study showed that prompt design plays a critical role in shaping output quality. Through collaboration between clinicians and engineers, we created more effective prompts and recommend that others use the CARE framework (Context, Action, Result, Example) to ensure the outputs are accurate, relevant, and actionable. Clinician involvement is essential not just for evaluating the LLM’s performance, but also for shaping its development. By participating in prompt engineering, clinicians move from passive users to co-developers. Their input ensures that LLMs in primary care are both technically functional and aligned with patient needs, professional values, and the realities of frontline practice.

  • Artificial Intelligence
  • Clinical Decision Support Systems
  • Natural Language Processing
  • Medical Informatics
  • Pilot Studies
  • Primary Health Care

In 2022, ChatGPT made artificial intelligence (AI) more accessible to health care clinicians by enabling them to engage with it using everyday language.1 These large language models (LLMs) have since allowed clinicians to learn from unstructured data and enhance communication with patients. Through this innovation, clinicians can use LLMs to document notes, respond to portal messages, summarize records, synthesize medical evidence, and create educational materials.2–4 These LLM outputs rated high for accuracy, readability, completeness, and empathy, at times exceeding scores for those generated by clinicians.5–7 Despite these benefits, clinicians remain ambivalent about using LLMs in practice. Some acknowledge their utility, while others are wary of their inconsistent performance, risks to patient safety, and limited transparency.8 In recent years, these use cases have been tested in isolated, demonstration projects. One study found that clinicians accepted and adopted the use of AI to respond to portal messages, an innovation that led to lower measures of burnout.9

The path to using LLMs in clinical settings began with key breakthroughs in model design. Generative Pre-trained Transformer (GPT)-3 marked a departure from earlier models and was important for several reasons. Before GPT-3, programmers had to write code, fine-tune models, and supply thousands to millions of training examples for each specific task. Because of its unique design and the billions of patterns it learned during training, GPT-3 could handle a wide-range of tasks, like a generalist entering a new community, identifying needs, and offering support across many areas. Instead of requiring task-specific training, this LLM could complete tasks with only a few examples, and sometimes none at all.10,11 This shift has led to the rise of prompting, which allows non-programmers to guide model outputs and make sure that they are accurate and useful.12 Prompts are essential because LLMs have gaps. Unlike clinicians and patients, they lack lived experience navigating the health care system and understanding what individual patients value most. They also do not consistently employ structured, rule-based reasoning so often used in clinical decision making. Well-crafted prompts guide the model’s reasoning, ensuring it approaches problems systematically. Ranging from simple instructions to complex frameworks, these prompts are best designed by clinicians, who can align outputs with clinical relevance and trustworthiness.

With prompting now central to LLMs, the role of subject matter experts has become increasingly important. These experts are essential for evaluating outputs and shaping effective prompts. Without experts to create the prompts, programmers, who are skilled in building and refining LLMs, may produce outputs that are inaccurate, irrelevant, or unsafe. Together, subject matter experts and programmers form a powerful dyad. Think of the programmers as architects who make sure that a house is structurally sound and that its plumbing and electricity are properly connected. The subject matter experts, then, are designers, selecting the right materials, ensuring that the space aligns with its intended purpose, and embedding meaning into every element. Both are needed to transform a well-built house into a functional and welcoming home.

The importance of this partnership became evident within our research team, which consisted of a family physician, health services researcher, and engineers. Together, we were building Primary Care (PC) Navigator, a multimodal AI tool that integrates an LLM with audio transcripts and facial-expression video recorded during encounters. The tool was pilot tested in a primary care clinic and was being used to develop 5A (assess, advise, agree, assist, and arrange) behavior change action plans that were printed out and provided to patients and clinicians at the end of their visits. This work was driven by data showing that only one-third of Americans receive counseling,13 and when counseling does occur, clinicians spend less than a minute.14 We built PC Navigator to fill this gap so that all patients receive adequate counseling tailored to their needs. The validity of plans generated by humans and PC Navigator, across a range of patients, is currently being studied.

Our pilot highlighted the critical role of prompt design in shaping LLM outputs. Our initial prompts lacked sufficient detail, resulting in outputs that were not clinically useful (Table 1). Clinicians noted that the responses were inconsistent in format, overly generic, failed to communicate clear goals, and included information unrelated to behavior change. After 8-10 iterations over several weeks, we created a new prompt to address these deficiencies (Table 1). The modified prompt defined 5A’s, directed the LLM to focus on health behaviors, and asked the responses to be grounded in the medical and social context. The prompt also specified the desired format, provided guidance on integrating emotional data from the encounter, and instructed the LLM to declare the ultimate goals and propose actionable next steps with detailed, realistic advice.

View this table:
  • View inline
  • View popup
Table 1. Original and Modified Prompts.

To help other primary care clinicians, learners, and researchers working with LLMs, we offer a practical approach for designing effective prompts. First, we recommend using the CARE (context, action, result, example) framework as a structured method for prompt creation (Table 1) and believe that prompt engineering skills will need to be incorporated into undergraduate and graduate training, particularly if the use of LLMs is shown to improve patient or learner outcomes.15 The context sets the stage by providing background information. The action specifies the task the LLM should perform. During this step, clinicians can employ a chain-of-thought approach that guides LLMs to generate intermediate steps before producing the final output. For instance, they might ask the LLM to detect emotional shifts during the visit, extract medical problems and social risk factors, uncover underlying beliefs and knowledge regarding the health behavior, clarify agreed-upon goals, outline the 5A components, and finally identify deeper barriers to change. The result articulates the desired outcome with pre-specified section headings. Finally, the example gives the model a clear reference for the ideal response. Including an example makes the prompt a one-shot prompt (where one “shot” refers to one example). Prompts can also be few-shot (multiple examples) or zero-shot (instructions only), depending on the LLM’s familiarity with the task, the amount of personalization required, and the computational resources available.

In conclusion, generating effective LLM prompts for primary care requires more than technical expertise. It also requires the active involvement of clinicians to ensure outputs resonate with patients and are medically appropriate. Without clinicians, we are concerned that the outputs risk being irrelevant to patients, sounding plausible but incorrect, or reinforcing existing biases. Clinicians bring a depth of knowledge that is not easily replicated, allowing them to discriminate between outputs that are inaccurate, dangerous, and confusing from those that are medically sound, clinically relevant, clear, and effective. They know what language will resonate with their patients and how to tailor responses to match cultural, professional, and emotional expectations. Their synergistic role in prompt engineering also ensures that outputs will be actionable, relevant, timely, and appropriate to different clinical settings, patient needs, and levels of risk. Hence, prompt engineering empowers clinicians not just to evaluate AI outputs, but to shape them. Using words rather than code, they are needed as co-developers who guide LLMs to produce content that aligns with real-world practice. Just as designers create spaces they would want to live and work in, clinicians create tools they would want to use in practice. Whenever LLMs are being developed for primary care, their involvement is mandatory so that the final product feels not like a generic solution developed for all clinical settings broadly but rather like it was something that was made specifically for primary care clinicians and patients.

Conflicts of Interest

Winston Liaw is a consultant for MedirAI and the American Academy of Family Physicians.

Corresponding Author

Winston Liaw, MD, MPH Department of Health Systems and Population Health Sciences, Tilman J. Fertitta Family College of Medicine, wliaw{at}central.uh.edu

This article was externally peer reviewed.

  • Received for publication August 18, 2025.
  • Accepted for publication October 27, 2025.

References

  1. ↵
    1. OpenAI
    (11 30, 2022) Introducing ChatGPT. . 2025-7-31. https://openai.com/index/chatgpt/.
  2. ↵
    1. Sallam M.
    (2023) ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. Healthcare 11(6):887, doi:10.3390/healthcare11060887, https://doi.org/10.3390/healthcare11060887. .
    OpenUrlCrossRefPubMed
  3. ↵
    1. Tang L.,
    2. Sun Z.,
    3. Idnay B..,
    4. et al.
    (2023) Evaluating large language models on medical evidence summarization. Npjdigitmed 6(1):158, doi:10.1038/s41746-023-00896-7, https://doi.org/10.1038/s41746-023-00896-7. .
    OpenUrlCrossRefPubMed
  4. ↵
    1. Guevara M.,
    2. Chen S.,
    3. Thomas S..,
    4. et al.
    (2024) Large language models to identify social determinants of health in electronic health records. Npjdigitmed 7(1):6, doi:10.1038/s41746-023-00970-0, https://doi.org/10.1038/s41746-023-00970-0. .
    OpenUrlCrossRef
  5. ↵
    1. Kaur A.,
    2. Budko A.,
    3. Liu K.,
    4. Eaton E.,
    5. Steitz B.D.,
    6. Johnson K.B.
    (2025) Automating responses to patient portal messages using generative AI. Appl Clin Inform 16(03):718–731, doi:10.1055/a-2565-9155, https://doi.org/10.1055/a-2565-9155. .
    OpenUrlCrossRef
  6. ↵
    1. Zaretsky J.,
    2. Kim J. M.,
    3. Baskharoun S..,
    4. et al.
    (2024) Generative artificial intelligence to transform inpatient discharge summaries to patient-friendly language and format. JAMA Netw Open 7(3):e240357, doi:10.1001/jamanetworkopen.2024.0357, https://doi.org/10.1001/jamanetworkopen.2024.0357. .
    OpenUrlCrossRefPubMed
  7. ↵
    1. Koh M. C. Y.,
    2. Ngiam J. N.,
    3. Oon J. E. L.,
    4. Lum L. H. W.,
    5. Smitasin N.,
    6. Archuleta S.
    (2025) Using ChatGPT for writing hospital inpatient discharge summaries – perspectives from an inpatient infectious diseases service. BMC Health Serv Res 25(1):221, doi:10.1186/s12913-025-12373-w, https://doi.org/10.1186/s12913-025-12373-w. .
    OpenUrlCrossRefPubMed
  8. ↵
    1. Hu D.,
    2. Guo Y.,
    3. Zhou Y.,
    4. Flores L.,
    5. Zheng K.
    (2025) A systematic review of early evidence on generative AI for drafting responses to patient messages. NpJ Health Syst 2:27, doi:10.1038/s44401-025-00032-5, https://doi.org/10.1038/s44401-025-00032-5. .
    OpenUrlCrossRef
  9. ↵
    1. Garcia P.,
    2. Ma S.P.,
    3. Shah S..,
    4. et al.
    (2024) Artificial intelligence–generated draft replies to patient inbox messages. JAMA Netw Open 7(3):e243201, doi:10.1001/jamanetworkopen.2024.3201, https://doi.org/10.1001/jamanetworkopen.2024.3201. .
    OpenUrlCrossRef
  10. ↵
    1. Reynolds L.,
    2. McDonell K.
    (2021) Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems (ACM), 1–7, doi:10.1145/3411763.3451760, https://doi.org/10.1145/3411763.3451760. . Prompt programming for large language models: beyond the few-shot paradigm.
    OpenUrlCrossRef
  11. ↵
    1. Brown T. B.,
    2. Mann B.,
    3. Ryder N..,
    4. et al.
    (7 22, 2020) Language models are few-shot learners. arXiv doi:10.48550/arXiv.2005.14165, https://doi.org/10.48550/arXiv.2005.14165. .
    OpenUrlCrossRef
  12. ↵
    1. Sivarajkumar S.,
    2. Kelley M.,
    3. Samolyk-Mazzanti A.,
    4. Visweswaran S.,
    5. Wang Y.
    (2024) An empirical evaluation of prompting strategies for large language models in zero-shot clinical natural language processing: algorithm development and validation study. JMIR Med Inform 12:e55318, doi:10.2196/55318, https://doi.org/10.2196/55318. .
    OpenUrlCrossRef
  13. ↵
    1. Ahmed N. U.,
    2. Delgado M.,
    3. Saxena A.
    (2017) Trends and disparities in the prevalence of physicians' counseling on exercise among the U.S. adult population, 2000–2010. Prev Med 99:1–6, doi:10.1016/j.ypmed.2017.01.015, https://doi.org/10.1016/j.ypmed.2017.01.015. .
    OpenUrlCrossRefPubMed
  14. ↵
    1. Podl T. R.,
    2. Goodwin M. A.,
    3. Kikano G. E.,
    4. Stange K. C.
    (1999) Direct observation of exercise counseling in community family practice. Am J Prev Med 17(3):207–210, doi:10.1016/S0749-3797(99)00074-4, https://doi.org/10.1016/S0749-3797(99)00074-4. .
    OpenUrlCrossRefPubMed
  15. ↵
    1. Juuzt AI
    (2025) The CARE Framework. . 2025-7-31. https://juuzt.ai/knowledge-base/prompt-frameworks/the-care-framework/.
PreviousNext
Back to top

In this issue

The Journal of the American Board of Family     Medicine: 39 (1)
The Journal of the American Board of Family Medicine
Vol. 39, Issue 1
1 Jul 2026
  • Table of Contents
  • Index by author
Print
Download PDF
Article Alerts
Sign In to Email Alerts with your Email Address
Email Article

Thank you for your interest in spreading the word on American Board of Family Medicine.

NOTE: We only request your email address so that the person you are recommending the page to knows that you wanted them to see it, and that it is not junk mail. We do not capture any email address.

Enter multiple addresses on separate lines or separate them with commas.
The Importance of Primary Care Subject Matter Experts: Output Quality in Large Language Models Prompt Engineering
(Your Name) has sent you a message from American Board of Family Medicine
(Your Name) thought you would like to see the American Board of Family Medicine web site.
CAPTCHA
This question is for testing whether or not you are a human visitor and to prevent automated spam submissions.
Citation Tools
The Importance of Primary Care Subject Matter Experts: Output Quality in Large Language Models Prompt Engineering
Winston Liaw, Quang Hung Bui, Omolola Adepoju, Hien Van Nguyen
The Journal of the American Board of Family Medicine Jun 2026, 39 (1) 162476; DOI: 10.3122/jabfm.2025.250326R1

Citation Manager Formats

  • BibTeX
  • Bookends
  • EasyBib
  • EndNote (tagged)
  • EndNote 8 (xml)
  • Medlars
  • Mendeley
  • Papers
  • RefWorks Tagged
  • Ref Manager
  • RIS
  • Zotero
Share
The Importance of Primary Care Subject Matter Experts: Output Quality in Large Language Models Prompt Engineering
Winston Liaw, Quang Hung Bui, Omolola Adepoju, Hien Van Nguyen
The Journal of the American Board of Family Medicine Jun 2026, 39 (1) 162476; DOI: 10.3122/jabfm.2025.250326R1
Twitter logo Facebook logo Mendeley logo
  • Tweet Widget
  • Facebook Like
  • Google Plus One

Jump to section

  • Article
    • Abstract
    • Conflicts of Interest
    • Corresponding Author
    • References
  • References
  • Info & Metrics
  • PDF

Related Articles

  • No related articles found.
  • PubMed
  • Google Scholar

Cited By...

  • No citing articles found.
  • Google Scholar

More in this TOC Section

  • “Hard Fork” for Family Medicine - Artificial Intelligence Will Change the Way We Experience Practice
  • From Pipeline to Practice: Supporting Family Physicians as a Maternal Health Solution
Show more Commentary

Similar Articles

Keywords

  • Artificial Intelligence
  • Clinical Decision Support Systems
  • Natural Language Processing
  • Medical Informatics
  • Pilot Studies
  • Primary Health Care

Navigate

  • Home
  • Current Issue
  • Past Issues

Authors & Reviewers

  • Info For Authors
  • Info For Reviewers
  • Submit A Manuscript/Review

Other Services

  • Get Email Alerts
  • Classifieds
  • Reprints and Permissions

Other Resources

  • Forms
  • Contact Us
  • ABFM News

© 2026 American Board of Family Medicine

Powered by HighWire