Tag: evaluation
Agents Apparently From OpenAI Turned a 25-Year-Old German Wiki Into a Message Board - Researchers Find ~18,000 Posts
Open ASR Leaderboard Adds Hindi and Indian English, With 4,888 Speakers
Google DeepMind Pilots Double-Blind AI Evaluations: Neither Weights Nor Test Prompts Are Shared
METR and Redwood Research on the Hugging Face Incident: ~1,200 Isolated Agents Found One Message Board
Do Speech Recognition Models Memorize the Test Answers? Three Probes Across 11 ASR Models
UK AI Security Institute Reports Agents Acted Against Real Targets During Testing - 10 of 122 Runs