Escrever uma avaliação PREreview

Towards Evaluating the Diagnostic Ability of LLMs

de Peter Sarvari e Zaid Al-fagih

Publicado: 12 de outubro de 2024
Servidor: Preprints.org
DOI: 10.20944/preprints202409.0688.v3

On average, one in ten patients die because of a diagnostic error and medical errors are the third largest cause of death in the US. While LLMs have been proposed to help doctors with diagnoses, no research results have been published on comparing the diagnostic ability of many popular LLMs on an openly accessible real-patient cohort. In thus study, we compare LLMs from Google, OpenAI, Meta, Mistral, Cohere and Anthropic using our previously published evaluation methodology and explore improving their accuracy with RAG.

Você pode escrever uma avaliação PREreview de Towards Evaluating the Diagnostic Ability of LLMs. Uma avaliação PREreview é uma avaliação de um preprint e pode variar de algumas frases a um parecer extenso, semelhante a um parecer de revisão por pares realizado por periódicos.

Antes de começar

Vamos pedir que você faça login com seu ORCID iD. Se você não tiver um iD, pode criar um.