AI delivers helpful heart health advice—but not every time
Artificial intelligence (AI) models continue to evolve. However, it remains unclear if they can provide heart-healthy diet and exercise recommendations that line up with established society guidelines.
To learn more, a team of researchers evaluated the effectiveness of four popular large language models (LLMs)—ChatGPT, Claude AI, DeepSeek AI and Google Gemini—when asked a series of questions about cardiovascular health. The group shared its findings in Cureus.[1]
“This study aimed to assess the appropriateness, comprehensibility, and clinical relevance of chatbot-generated recommendations to evaluate their potential as reliable adjunct tools in preventive cardiology and to ensure alignment with established evidence-based guidelines from leading cardiovascular organizations,” wrote first author Tagbo C. Nduka, MD, a researcher with Texas A&M Medicine, and colleagues. “Additionally, it seeks to identify possible instances of misinformation and determine whether these AI tools align with evidence-based recommendations provided by reputable cardiovascular organizations.”
Nduka et al. asked the four LLMs a total of 15 questions. Each response was then categorized as appropriate, appropriate but insufficient, partially inappropriate or entirely inappropriate based on how well they comply with heart health recommendations from the American College of Cardiology (ACC), American Heart Association (AHA) and European Society of Cardiology (ESC).
Overall, 90% of responses from ChatGPT, Claude AI and DeepSeek AI appropriately lined up with society guidelines. The same was true for 80% of responses from Google Gemini.
“Our findings indicated that expert medical counsel cannot be substituted by language models, notwithstanding their ability to provide readily accessible and fundamentally appropriate health information,” the authors wrote. “Further research may be necessary to fully understand the capabilities of these language models.”
One big takeaway from this analysis was the fact that LLMs seemed to consistently prefer referencing some industry guidelines more than others.
“The responses produced by the AI system demonstrated a stronger inclination for the recommendations of the AHA/ACC relative to those offered by the ESC,” the authors added. “This compelling discovery may arise from the inherent biases of the AI-powered tools or their developers. It is crucial to meticulously assess the recommendations of these LLMs to reduce biased responses.”
Click here to read the full analysis.
