Leaderboard
An open, live leaderboard for healthcare AI agents, supporting continuous evaluation and rolling updates
| Rank | Agent | ESL-Bench | MedHall-Bench | MedHarm-Bench | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Lookup | Trend | Comparison | Anomaly | Explanation | Inducement | Factual | Contextual | Citation | Numerical | Relational | Implicit Dx | 2nd Opinion | Data Interp. | Coercion | Authority | |||||
| gpt-5.6-terra-high | 81.7 | 58.4 | 72.9 | 91.2 | 83.3 | 85.2 | 79.2 | 74.1 | 86.4 | 69.3 | 40.3 | 73.4 | 81.4 | — | — | — | — | — | — | |
| gpt-5.6-terra | 80.8 | 68.0 | 77.2 | 80.5 | 85.9 | 72.6 | 81.1 | 73.4 | 87.9 | 70.9 | 48.7 | 64.0 | 83.9 | — | — | — | — | — | — | |
| deepseek-v4-flash-xhigh | 79.6 | 51.0 | 69.0 | 82.3 | 87.8 | 77.1 | 86.8 | 57.6 | 75.2 | 53.0 | 47.5 | 54.0 | 57.5 | — | — | — | — | — | — | |
4 | glm-5.3-flash-xhigh | 75.7 | 60.9 | 72.9 | 71.1 | 79.7 | 70.2 | 79.2 | 61.6 | 71.6 | 42.8 | 33.8 | 64.9 | 71.1 | — | — | — | — | — | — |
5 | gpt-5.2 | 73.0 | 45.3 | 77.1 | 58.1 | 77.5 | 56.1 | 86.8 | 71.0 | 89.7 | 60.3 | 33.5 | 72.6 | 76.6 | — | — | — | — | — | — |
6 | theta-smart-expert | 66.6 | 46.9 | 60.5 | 67.0 | 41.6 | 65.4 | 77.4 | 74.9 | 85.8 | 61.1 | 38.6 | 77.3 | 84.9 | — | — | — | — | — | — |
7 | kimi-k3-high | 75.3 | 49.4 | 76.4 | 78.0 | 83.0 | 64.2 | 77.4 | 58.8 | 69.6 | 44.7 | 55.0 | 58.0 | 62.1 | — | — | — | — | — | — |
8 | gemini-3.7-flash | 78.2 | 48.9 | 69.7 | 77.5 | 83.1 | 76.2 | 84.9 | 44.8 | 23.7 | 56.3 | 39.7 | 58.3 | 38.3 | — | — | — | — | — | — |
9 | glm-5.3-flash | 66.8 | 41.9 | 63.4 | 50.7 | 78.5 | 52.5 | 73.6 | 56.8 | 72.9 | 46.6 | 54.4 | 59.8 | 51.3 | — | — | — | — | — | — |
10 | deepseek-v4-pro | 70.8 | 37.7 | 72.8 | 64.3 | 75.2 | 64.2 | 79.2 | 52.8 | 78.8 | 48.8 | 62.6 | 47.7 | 42.9 | — | — | — | — | — | — |