Close Menu
TechurzTechurz
    What's Hot

    Bot-detection startup Spur nabs $200M from Insight

    July 28, 2026

    MCP startup Runlayer accuses Rippling of stealing its product idea

    July 28, 2026

    Ozlo’s Sleepbuds 2 build on Bose’s sleep earbud legacy

    July 28, 2026
    X (Twitter) Pinterest YouTube LinkedIn WhatsApp
    Tech Pulse
    • Bot-detection startup Spur nabs $200M from Insight
    • MCP startup Runlayer accuses Rippling of stealing its product idea
    • Ozlo’s Sleepbuds 2 build on Bose’s sleep earbud legacy
    • Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing
    • Antares raises $470M to build nuclear reactors for the US military
    X (Twitter) Pinterest YouTube LinkedIn WhatsApp
    TechurzTechurz
    • Home
    • Tech Pulse
    • Future Tech
    • AI Systems
    • Cyber Reality
    • Disruption Lab
    • Signals
    TechurzTechurz
    Home - Apps - Every AI model is flunking medicine – and LMArena proposes a fix
    Apps

    Every AI model is flunking medicine – and LMArena proposes a fix

    TechurzBy TechurzAugust 19, 2025Updated:May 11, 2026No Comments4 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Every AI model is flunking medicine - and LMArena proposes a fix
    Share
    Facebook Twitter LinkedIn Pinterest Email


    johan63/iStock/Getty Images Plus via Getty Images

    Table of contents
    1 ZDNET’s key takeaways
    2 A knowledge gap in medicine
    3 Expanding the benchmark
    4 Where could they go wrong?

    ZDNET’s key takeaways

    • AI frontier models fail to provide safe and accurate output on medical topics.
    • LMArena and DataTecnica aim to ‘rigorously’ test LLMs’ medical knowledge.
    • It’s not clear how agents and medicine-specific LLMs will be measured.

    Get more in-depth ZDNET tech coverage: Add us as a preferred Google source on Chrome and Chromium browsers.

    Despite the numerous AI advances in medicine cited throughout scholarly literature, all generative AI programs fail to produce output that is both safe and accurate when dealing with medical topics, according to a new report by benchmark firm LMArena. 

    The finding is especially concerning given that people are going to bots such as ChatGPT for medical answers, and research shows that people trust AI’s medical advice over the advice of doctors, even when it’s wrong.

    Also: Patients trust AI’s medical advice over doctors – even when it’s wrong, study finds

    The new study, comparing OpenAI’s GPT-5 with numerous models from Google, Anthropic, and Meta, finds that “performance in real-world biomedical research remains far from adequate.” 

    (Disclosure: Ziff Davis, ZDNET’s parent company, filed an April 2025 lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)

    A knowledge gap in medicine

    “No current model reliably meets the reasoning and domain-specific knowledge demands of biomedical scientists,” according to the LMArena team.

    The report concludes that current models are simply too lax and too fuzzy to meet the standards of medicine:

    “This fundamental gap highlights the growing mismatch between general AI capabilities and the needs of specialized scientific communities. Biomedical researchers work at the intersection of complex, evolving knowledge and real-world impact. They don’t need models that ‘sound’ correct; they need tools that help uncover insights, reduce error, and accelerate the pace of discovery.”

    LMArena + DataTecnica

    The study echoes findings from other benchmark tests related to medicine. For example, in May, OpenAI unveiled HealthBench, a suite of text prompts concerning medical situations and conditions that could reasonably be submitted to a chatbot by a person seeking medical advice. That study found that the best accuracy score, by OpenAI’s o3 large language model, 0.598, left ample room for improvement on the benchmark. 

    Also: OpenAI’s HealthBench shows AI’s medical advice is improving – but who will listen?

    Expanding the benchmark

    To address the gap between AI models and medicine, LMArena has teamed with startup DataTecnica, which earlier this year unveiled a benchmark suite of tests for Gen AI called CARDBiomedBench, a question-and-answer benchmark for evaluating LLMs in biomedical research.

    Together, LMArena and DataTecnica plan to expand what’s called BiomedArena, a leaderboard that lets people compare AI models side by side and vote on which ones perform the best.

    Also: Meta’s Llama 4 ‘herd’ controversy and AI contamination, explained

    BiomedArena is meant to be specific to medical research, rather than very general questions, unlike general-purpose leaderboards.

    The BiomedArena work is already used by scientists at the Intramural Research Program of the US National Institutes of Health, they note, “where scientists pursue high-risk, high-reward projects that are often beyond the scope of traditional academic research due to their scale, complexity, or resource demands.”

    The BiomedArena work, according to the LMArena team, will “focus on tasks and evaluation strategies grounded in the day-to-day realities of biomedical discovery — from interpreting experimental data and literature to assisting in hypothesis generation and clinical translation.”

    Also: You can track the top AI image generators via this new leaderboard – and vote for your favorite too

    As ZDNET’s Webb Wright reported in June, LMArena.ai ranks AI models. The website was originally founded as a research initiative through UC Berkeley under the name Chatbot Arena and has since become a full-fledged platform, with financial support from UC Berkeley, a16z, Sequoia Capital, and others.

    Where could they go wrong?

    Two big questions loom for this new benchmark effort.

    First, studies with doctors have shown that gen AI’s usefulness expands dramatically when AI models are hooked up to databases of “gold standard” medical information, with dedicated large language models (LLMs) able to outperform the top frontier models just by tapping into information. 

    Also: Hooking up generative AI to medical data improved usefulness for doctors

    From today’s announcement, it’s not clear how LMArena and DataTecnica plan to address that aspect of AI models, which really is a kind of agentic capability — the ability to tap into resources. Without measuring how AI models use external resources, the benchmark could have limited utility.

    Second, numerous medicine-specific LLMs are being developed all the time, including Google’s “MedPaLM” program developed two years ago. It’s not clear if the BiomedArena work will take into account these dedicated medicine LLMs. The work so far has tested only general frontier models. 

    Also: Google’s MedPaLM emphasizes human clinicians in medical AI

    That’s a perfectly valid choice on the part of LMArena and DataTecnica, but it does leave out a whole lot of important effort.

    fix flunking LMArena medicine model proposes
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleElectronic Health Record Giant Epic Rolling Out New AI Tools
    Next Article Samsung will give you a free 65-inch TV right now – here’s how to get one
    Techurz
    • Website

    Related Posts

    Opinion

    Runway launches AI model router as generative media gets crowded

    July 23, 2026
    Opinion

    Applied Computing wants to give oil and gas operators an AI model for the entire plant

    July 16, 2026
    Opinion

    Backed by $60M in funding, Oak steps out of stealth to fix the identity mess that AI agents are making worse

    July 15, 2026
    Add A Comment
    Latest Tech Pulse

    College social app Fizz expands into grocery delivery

    September 3, 20252,290

    12 Father’s Day E-Card Sites That Are Actually Good

    June 4, 202523

    SolarSquare in talks to raise up to $60M as India’s rooftop solar market draws major VC interest

    May 23, 202622
    Stay In Touch
    • YouTube
    • WhatsApp
    • Twitter
    • Pinterest
    • LinkedIn

    Techurz helps readers stay ahead of digital change with clear, practical, future focused technology intelligence written today,searched tomorrow.

    X (Twitter) Pinterest YouTube LinkedIn WhatsApp
    Company
    • About Us
    • Contact Us
    • Our Authors / Editorial Team
    • Write For Us
    • Advertise
    Policy
    • Editorial Policy
    • Privacy Policy
    • Terms and Conditions
    • Affiliate Disclosure
    • Cookie Policy
    • Disclaimer
    • DMCA
    Explore
    • AI Systems
    • Cyber Reality
    • Future Tech
    • Disruption Lab
    • Signals
    • Tech Pulse
    • Sitemap

    Join the Techurz Brief

    The future does not arrive suddenly.
    Stay ahead with fast, sharp tech signals.

    Type above and press Enter to search. Press Esc to cancel.