#17. MedHELM: Avaliando LLMs para Tarefas Médicas

Failed to add items

Sorry, we are unable to add the item because your shopping cart is already at capacity.

Add to basket failed.

Please try again later

Add to Wish List failed.

Please try again later

Remove from Wish List failed.

Please try again later

Follow podcast failed

Unfollow podcast failed

#17. MedHELM: Avaliando LLMs para Tarefas Médicas

Listen for free

View show details

About this listen

O estudo MedHELM apresenta uma estrutura de avaliação abrangente para Grandes Modelos de Linguagem (LLMs) em tarefas médicas, indo além dos testes de licenciamento tradicionais. Ele introduz uma taxonomia validada por clínicos com 5 categorias, 22 subcategorias e 121 tarefas, juntamente com um conjunto de 35 benchmarks, incluindo 18 novos. A avaliação de nove LLMs de ponta revelou que modelos de raciocínio, como DeepSeek R1 e o3-mini, demonstram desempenho superior, embora Claude 3.5 Sonnet ofereça um equilíbrio de custo-benefício. O sistema de "júri de LLMs" proposto demonstrou maior concordância com avaliações clínicas do que métodos automatizados convencionais. Essas descobertas destacam a importância de avaliações focadas em tarefas do mundo real para a implementação segura de LLMs na saúde.

What listeners say about #17. MedHELM: Avaliando LLMs para Tarefas Médicas

Average Customer Ratings

Reviews - Please select the tabs below to change the source of reviews.

Audible.com.au reviews

Amazon Reviews

No Reviews are Available

Report a review on Amazon

Audiobook Categories

More to Explore

GETTING STARTED

#17. MedHELM: Avaliando LLMs para Tarefas Médicas

Failed to add items

Add to basket failed.

Add to Wish List failed.

Remove from Wish List failed.

Follow podcast failed

Unfollow podcast failed

#17. MedHELM: Avaliando LLMs para Tarefas Médicas

About this listen

What listeners say about #17. MedHELM: Avaliando LLMs para Tarefas Médicas

Reviews - Please select the tabs below to change the source of reviews.

Audible.com.au reviews

Amazon Reviews