Results on January 10, 2025
Multiple Choice Track
Model | Submission Time (GMT) | Original | NOTA |
---|
Claude-3.5 Sonnet | 2025-01-11 03:00:00 | 75.0 | 55.0 |
Gemini 1.5 Flash | 2025-01-11 03:00:00 | 70.0 | 50.0 |
GPT-3.5 Turbo + Google Custom Search | 2025-01-11 03:00:00 | 65.0 | 55.0 |
GPT-4o + Google Custom Search | 2025-01-11 03:00:00 | 65.0 | 55.0 |
GPT-4o | 2025-01-11 03:00:00 | 65.0 | 45.0 |
Gemini 1.5 Flash + Google Custom Search | 2025-01-11 03:00:00 | 60.0 | 50.0 |
Claude-3.5 Haiku | 2025-01-11 03:00:00 | 60.0 | 45.0 |
GPT-3.5 Turbo | 2025-01-11 03:00:00 | 60.0 | 25.0 |
Claude-3.5 Haiku + Google Custom Search | 2025-01-11 03:00:00 | 50.0 | 50.0 |
Claude-3.5 Sonnet + Google Custom Search | 2025-01-11 03:00:00 | 50.0 | 45.0 |
Llama3.1-405B-Instruct + Google Custom Search | 2025-01-11 03:00:00 | 45.0 | 40.0 |
Llama3.1-405B-Instruct | 2025-01-11 03:00:00 | 35.0 | 35.0 |
Generation Track
Model | Submission Time (GMT) | EM | F1 |
---|
GPT-4o | 2025-01-11 03:00:00 | 40.0 | 49.2 |
Gemini 1.5 Flash | 2025-01-11 03:00:00 | 30.0 | 42.2 |
Llama3.1-405B-Instruct | 2025-01-11 03:00:00 | 25.0 | 41.3 |
Llama3.1-405B-Instruct + Google Custom Search | 2025-01-11 03:00:00 | 25.0 | 39.8 |
GPT-3.5 Turbo + Google Custom Search | 2025-01-11 03:00:00 | 25.0 | 38.5 |
GPT-4o + Google Custom Search | 2025-01-11 03:00:00 | 25.0 | 36.9 |
GPT-3.5 Turbo | 2025-01-11 03:00:00 | 20.0 | 36.3 |
Gemini 1.5 Flash + Google Custom Search | 2025-01-11 03:00:00 | 15.0 | 23.1 |
Claude-3.5 Haiku | 2025-01-11 03:00:00 | 0.0 | 15.2 |
Claude-3.5 Haiku + Google Custom Search | 2025-01-11 03:00:00 | 0.0 | 14.5 |
Claude-3.5 Sonnet | 2025-01-11 03:00:00 | 0.0 | 12.5 |
Claude-3.5 Sonnet + Google Custom Search | 2025-01-11 03:00:00 | 0.0 | 6.2 |