Abstract
Large language models (LLMs) offer new opportunities to support primary care by generating documentation, summarizing records, answering messages, and providing education. While LLMs show promise, their effectiveness depends on how well their outputs align with real-world clinical needs. This commentary describes one team’s development of Primary Care (PC) Navigator, a multimodal artificial intelligence (AI) tool that combines audio and video from clinical encounters with an LLM to generate behavior change plans. Our pilot study showed that prompt design plays a critical role in shaping output quality. Through collaboration between clinicians and engineers, we created more effective prompts and recommend that others use the CARE framework (Context, Action, Result, Example) to ensure the outputs are accurate, relevant, and actionable. Clinician involvement is essential not just for evaluating the LLM’s performance, but also for shaping its development. By participating in prompt engineering, clinicians move from passive users to co-developers. Their input ensures that LLMs in primary care are both technically functional and aligned with patient needs, professional values, and the realities of frontline practice.
- Artificial Intelligence
- Clinical Decision Support Systems
- Natural Language Processing
- Medical Informatics
- Pilot Studies
- Primary Health Care
In 2022, ChatGPT made artificial intelligence (AI) more accessible to health care clinicians by enabling them to engage with it using everyday language.1 These large language models (LLMs) have since allowed clinicians to learn from unstructured data and enhance communication with patients. Through this innovation, clinicians can use LLMs to document notes, respond to portal messages, summarize records, synthesize medical evidence, and create educational materials.2–4 These LLM outputs rated high for accuracy, readability, completeness, and empathy, at times exceeding scores for those generated by clinicians.5–7 Despite these benefits, clinicians remain ambivalent about using LLMs in practice. Some acknowledge their utility, while others are wary of their inconsistent performance, risks to patient safety, and limited transparency.8 In recent years, these use cases have been tested in isolated, demonstration projects. One study found that clinicians accepted and adopted the use of AI to respond to portal messages, an innovation that led to lower measures of burnout.9
The path to using LLMs in clinical settings began with key breakthroughs in model design. Generative Pre-trained Transformer (GPT)-3 marked a departure from earlier models and was important for several reasons. Before GPT-3, programmers had to write code, fine-tune models, and supply thousands to millions of training examples for each specific task. Because of its unique design and the billions of patterns it learned during training, GPT-3 could handle a wide-range of tasks, like a generalist entering a new community, identifying needs, and offering support across many areas. Instead of requiring task-specific training, this LLM could complete tasks with only a few examples, and sometimes none at all.10,11 This shift has led to the rise of prompting, which allows non-programmers to guide model outputs and make sure that they are accurate and useful.12 Prompts are essential because LLMs have gaps. Unlike clinicians and patients, they lack lived experience navigating the health care system and understanding what individual patients value most. They also do not consistently employ structured, rule-based reasoning so often used in clinical decision making. Well-crafted prompts guide the model’s reasoning, ensuring it approaches problems systematically. Ranging from simple instructions to complex frameworks, these prompts are best designed by clinicians, who can align outputs with clinical relevance and trustworthiness.
With prompting now central to LLMs, the role of subject matter experts has become increasingly important. These experts are essential for evaluating outputs and shaping effective prompts. Without experts to create the prompts, programmers, who are skilled in building and refining LLMs, may produce outputs that are inaccurate, irrelevant, or unsafe. Together, subject matter experts and programmers form a powerful dyad. Think of the programmers as architects who make sure that a house is structurally sound and that its plumbing and electricity are properly connected. The subject matter experts, then, are designers, selecting the right materials, ensuring that the space aligns with its intended purpose, and embedding meaning into every element. Both are needed to transform a well-built house into a functional and welcoming home.
The importance of this partnership became evident within our research team, which consisted of a family physician, health services researcher, and engineers. Together, we were building Primary Care (PC) Navigator, a multimodal AI tool that integrates an LLM with audio transcripts and facial-expression video recorded during encounters. The tool was pilot tested in a primary care clinic and was being used to develop 5A (assess, advise, agree, assist, and arrange) behavior change action plans that were printed out and provided to patients and clinicians at the end of their visits. This work was driven by data showing that only one-third of Americans receive counseling,13 and when counseling does occur, clinicians spend less than a minute.14 We built PC Navigator to fill this gap so that all patients receive adequate counseling tailored to their needs. The validity of plans generated by humans and PC Navigator, across a range of patients, is currently being studied.
Our pilot highlighted the critical role of prompt design in shaping LLM outputs. Our initial prompts lacked sufficient detail, resulting in outputs that were not clinically useful (Table 1). Clinicians noted that the responses were inconsistent in format, overly generic, failed to communicate clear goals, and included information unrelated to behavior change. After 8-10 iterations over several weeks, we created a new prompt to address these deficiencies (Table 1). The modified prompt defined 5A’s, directed the LLM to focus on health behaviors, and asked the responses to be grounded in the medical and social context. The prompt also specified the desired format, provided guidance on integrating emotional data from the encounter, and instructed the LLM to declare the ultimate goals and propose actionable next steps with detailed, realistic advice.
To help other primary care clinicians, learners, and researchers working with LLMs, we offer a practical approach for designing effective prompts. First, we recommend using the CARE (context, action, result, example) framework as a structured method for prompt creation (Table 1) and believe that prompt engineering skills will need to be incorporated into undergraduate and graduate training, particularly if the use of LLMs is shown to improve patient or learner outcomes.15 The context sets the stage by providing background information. The action specifies the task the LLM should perform. During this step, clinicians can employ a chain-of-thought approach that guides LLMs to generate intermediate steps before producing the final output. For instance, they might ask the LLM to detect emotional shifts during the visit, extract medical problems and social risk factors, uncover underlying beliefs and knowledge regarding the health behavior, clarify agreed-upon goals, outline the 5A components, and finally identify deeper barriers to change. The result articulates the desired outcome with pre-specified section headings. Finally, the example gives the model a clear reference for the ideal response. Including an example makes the prompt a one-shot prompt (where one “shot” refers to one example). Prompts can also be few-shot (multiple examples) or zero-shot (instructions only), depending on the LLM’s familiarity with the task, the amount of personalization required, and the computational resources available.
In conclusion, generating effective LLM prompts for primary care requires more than technical expertise. It also requires the active involvement of clinicians to ensure outputs resonate with patients and are medically appropriate. Without clinicians, we are concerned that the outputs risk being irrelevant to patients, sounding plausible but incorrect, or reinforcing existing biases. Clinicians bring a depth of knowledge that is not easily replicated, allowing them to discriminate between outputs that are inaccurate, dangerous, and confusing from those that are medically sound, clinically relevant, clear, and effective. They know what language will resonate with their patients and how to tailor responses to match cultural, professional, and emotional expectations. Their synergistic role in prompt engineering also ensures that outputs will be actionable, relevant, timely, and appropriate to different clinical settings, patient needs, and levels of risk. Hence, prompt engineering empowers clinicians not just to evaluate AI outputs, but to shape them. Using words rather than code, they are needed as co-developers who guide LLMs to produce content that aligns with real-world practice. Just as designers create spaces they would want to live and work in, clinicians create tools they would want to use in practice. Whenever LLMs are being developed for primary care, their involvement is mandatory so that the final product feels not like a generic solution developed for all clinical settings broadly but rather like it was something that was made specifically for primary care clinicians and patients.
Conflicts of Interest
Winston Liaw is a consultant for MedirAI and the American Academy of Family Physicians.
Corresponding Author
Winston Liaw, MD, MPH Department of Health Systems and Population Health Sciences, Tilman J. Fertitta Family College of Medicine, wliaw{at}central.uh.edu
This article was externally peer reviewed.
- Received for publication August 18, 2025.
- Accepted for publication October 27, 2025.






