DR. INFO Team Updates HealthBench Results for Clinical AI Assistant
Researchers revised their DR. INFO evaluation on July 27, 2026, reporting a 0.68 score on HealthBench Hard and 0.72 in a separate 100-case comparison with two other retrieval-based clinical assistants. The results are author-reported benchmark scores for an integrated retrieval system, not evidence of clinical safety, patient outcomes, or a standalone model advantage.
The team behind DR. INFO revised its HealthBench evaluation on July 27, 2026, adding updated scores and comparisons with the GPT-5 model family. The authors report that their agentic retrieval-augmented clinical assistant scored 0.68 on HealthBench Hard, a 1,000-conversation subset designed to challenge current systems.
What was evaluated
DR. INFO is an integrated system rather than a standalone language model. It decomposes clinical questions, retrieves material from curated medical sources and peer-reviewed literature, and synthesizes cited responses. The paper evaluates that full workflow on OpenAI's rubric-based HealthBench dataset.
The revised paper reports a 0.68 HealthBench Hard score for DR. INFO. It lists GPT-5-family scores between 0.40 and 0.46, Grok 3 at 0.23, Gemini 2.5 Pro at 0.19, and Claude 3.7 Sonnet at 0.02. In a separate 100-case comparison with other retrieval-based clinical assistants, the authors report 0.72 for DR. INFO, 0.49 for OpenEvidence, and 0.48 for DoxGPT.
Those figures are not clean model-to-model comparisons. The authors' companion Cureus paper says DR. INFO combines an underlying model with retrieval and workflow logic, while the general-purpose comparators were evaluated without the same clinical retrieval layer. The measured difference therefore reflects the combined architecture, data access, and orchestration—not just the base model.
What the benchmark can and cannot show
HealthBench uses realistic, open-ended health conversations and physician-written rubrics. OpenAI says the broader benchmark contains 5,000 conversations and that HealthBench Hard selects 1,000 particularly difficult examples. Responses are graded against conversation-specific criteria using a model-based evaluator.
The reported scores show how DR. INFO performed on that benchmark at the time of evaluation. They do not establish diagnostic accuracy in clinical deployment, safety across patient populations, or improved patient outcomes. The companion pilot study involved 29 physicians and medical students, relied largely on perceived usefulness, and called for larger controlled studies with independent accuracy verification.
For data and AI teams, the practical takeaway is to compare systems at the same layer. Retrieval, orchestration, and source curation can materially change benchmark performance, but a system-level score should not be marketed as proof that one underlying model is medically superior.
Key Points
- 1The revised paper reports a 0.68 score for DR. INFO on the 1,000-case HealthBench Hard subset.
- 2A separate 100-case comparison reports 0.72 for DR. INFO versus 0.49 for OpenEvidence and 0.48 for DoxGPT.
- 3The comparisons combine model, retrieval, and orchestration effects and do not establish clinical safety or patient benefit.
Scoring Rationale
The update is useful evidence about system-level evaluation of a clinical RAG assistant, but the results are author-reported, not independently replicated, and do not establish clinical outcomes.
Sources
Primary source and supporting public references used for this report.
Practice with real Health & Insurance data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all Health & Insurance problems
